Skip to main content
CareerApp

Research Engineer, Post-Training Inference

Together AI

San Francisco, CA · full time

$200,000 – $290,000

Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role involves building and optimizing systems that allow developers to fine-tune and deploy open-source AI models efficiently, working across the full pipeline from training through production inference. You'll collaborate on core infrastructure for customization services, focusing on inference optimization and integration between post-training and serving platforms.

Our summary, not Together AI’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 2+ years building and deploying ML-based services in production
  • Hands-on experience with modern inference engines such as SGLang, vLLM, or TensorRT-LLM
  • Familiarity with modern fine-tuning methods for LLMs and AI models
  • Strong software engineering background in Python or Go
  • Awareness of current advances and trends in machine learning

Nice to have

  • Experience serving low-precision models or managing multiple LoRA adapters in one instance
  • Experience optimizing RL training workloads
  • Experience developing CUDA, Triton, or CuTE DSL kernels for inference
  • Experience building large-scale, high-load production systems
  • Contributions to or maintenance of open-source ML projects
  • Experience managing ML workloads on Kubernetes

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.