# Distributed LLM Inference Engineer

Hiring organization: [Anyscale](https://career.thegoodapps.co/organizations/anyscale)

Canonical page: https://career.thegoodapps.co/jobs/82ff5dda-5e04-4607-a41e-6b3c06023890

Listed on Anyscale's own careers site. Applications go to them directly.

- Location: San Francisco, CA
- Remote: yes
- Salary: 170000 – 245000 USD per year

## Summary

Anyscale is hiring an engineer to optimize large-scale LLM inference systems, building infrastructure that developers can use to run machine learning models efficiently from laptop to cluster. This role suits someone with distributed systems expertise who wants to work on high-performance AI infrastructure and contribute to open-source projects like Ray and vLLM.

_Our summary, not Anyscale's wording._

## Skills named

CUDA, PyTorch, TensorFlow, V-Ray

## Required

- Familiarity with running ML inference at large scale with high throughput and low latency
- Familiarity with deep learning and deep learning frameworks
- Solid understanding of distributed systems and ML inference challenges

## Nice to have

- ML Systems knowledge
- Experience using Ray
- Work closely with community on LLM engines like vLLM and TensorRT-LLM
- Contributions to deep learning frameworks like PyTorch or TensorFlow
- Contributions to deep learning compilers like Triton, TVM, or MLIR
- Prior experience working on GPUs and CUDA

Apply on Anyscale's site: https://jobs.ashbyhq.com/anyscale/1cf38233-8aa0-47f8-9d85-65ce27bc3047/application
