$220,000 – $280,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Staff-level machine learning engineer role focused on optimizing inference for voice models like speech-to-text and text-to-speech on Together AI's platform. Ideal for someone with deep expertise in LLM serving engines and GPU optimization who wants to shape the technical direction of a foundational voice AI system.
Our summary, not Together AI’s wording. The full posting is on their site.
What they ask for
Required
- 8+ years of ML engineering experience with focus on model serving or inference optimization at production scale
- Deep expertise in LLM serving engines (vLLM, SGLang, TensorRT-LLM, or equivalent) including engine internals modification
- Expert-level Python and PyTorch with strong GPU optimization knowledge (CUDA, memory hierarchies, profiling)
- Proven system design judgment at scale
- Strong technical leadership with high autonomy
- Product intuition for developer tooling
- Ability to move fast in ambiguous, early-stage environments
- Foundation in speech and audio ML (ASR/TTS architectures, audio signal processing)
Nice to have
- Familiarity with audio codecs and tokenization schemes (SNAC, Encodec, DAC)
- Experience training or fine-tuning speech models at scale
- Bachelor's or Master's in Computer Science, Electrical Engineering, or related field