Skip to main content
CareerApp

Senior Backend Engineer, Inference Platform

Together AI

San Francisco, CA · full time · Senior

$160,000 – $250,000

Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role involves building and optimizing the core infrastructure that routes and balances inference requests across thousands of GPUs, working with cutting-edge AI hardware to make language models faster and more efficient at global scale. You'll partner with ML researchers to productionize new models while contributing to the open source tools that power the industry.

Our summary, not Together AI’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years building large-scale distributed systems and API microservices
  • Strong understanding of OS concepts: multi-threading, memory management, networking, storage performance
  • Expert-level programming in Rust, Go, Python, or TypeScript
  • Knowledge of modern LLMs and how they are served in production

Nice to have

  • Experience with the open source inference ecosystem (SGLang, vLLM, NVIDIA Dynamo)
  • Kubernetes or container orchestration experience
  • Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies (InfiniBand, NVLink, MPI)
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.