Skip to main content
CareerApp

Distributed LLM Inference Engineer

Anyscale

San Francisco, CA · Remote

$170,000 – $245,000

Listed on Anyscale’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Anyscale is hiring an engineer to optimize large-scale LLM inference systems, building infrastructure that developers can use to run machine learning models efficiently from laptop to cluster. This role suits someone with distributed systems expertise who wants to work on high-performance AI infrastructure and contribute to open-source projects like Ray and vLLM.

Our summary, not Anyscale’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Familiarity with running ML inference at large scale with high throughput and low latency
  • Familiarity with deep learning and deep learning frameworks
  • Solid understanding of distributed systems and ML inference challenges

Nice to have

  • ML Systems knowledge
  • Experience using Ray
  • Work closely with community on LLM engines like vLLM and TensorRT-LLM
  • Contributions to deep learning frameworks like PyTorch or TensorFlow
  • Contributions to deep learning compilers like Triton, TVM, or MLIR
  • Prior experience working on GPUs and CUDA

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.