Skip to main content
CareerApp

Machine Learning Engineer - Inference

Together AI

San Francisco, CA · full time · Mid

$160,000 – $230,000

Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

A Machine Learning Engineer role focused on building and optimizing inference systems for large language models at scale. This suits engineers with strong systems programming skills who want to work on high-performance AI infrastructure and collaborate closely with researchers.

Our summary, not Together AI’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 3+ years building production-quality, high-performance code
  • Proficiency with Python and PyTorch
  • Experience building high-performance libraries and tooling
  • Strong understanding of operating systems concepts including multi-threading, memory management, networking, and performance at scale

Nice to have

  • Knowledge of AI inference frameworks like TGI, vLLM, TensorRT-LLM, or Optimum
  • Knowledge of inference optimization techniques such as speculative decoding
  • CUDA or Triton programming experience
  • Familiarity with Rust, Cython, or compilers

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.