Skip to main content
CareerApp

Forward Deployed Engineer (Inference & Post-Training)

Together AI

San Francisco, CA · full time · Senior

$270,000 – $300,000

Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

A technical specialist working directly with large customers to optimize AI model inference and fine-tuning on production systems. This role suits engineers with deep expertise in LLM deployment and performance tuning who want to work closely with strategic customers while influencing product direction.

Our summary, not Together AI’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years in a technical role
  • Expert-level hands-on experience with inference engines like vLLM, TensorRT-LLM, or SGLang
  • Deep knowledge of KV cache tuning, speculative decoding, tensor parallelism, and quantization
  • Hands-on experience with fine-tuning pipelines including LoRA, SFT, DPO, RLHF, and GRPO
  • Strong Python skills
  • Broad knowledge of open-source LLM landscape

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.