$200,000 – $260,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role focuses on optimizing how voice AI models are served in production, tackling unique challenges like real-time latency and streaming audio. It suits experienced ML engineers comfortable diving deep into inference engines and GPU optimization who want to shape infrastructure as the industry moves toward end-to-end speech systems.
Our summary, not Together AI’s wording. The full posting is on their site.
What they ask for
Required
- 5+ years ML engineering experience
- Model serving or inference optimization focus
- Hands-on experience with LLM serving engines like vLLM, SGLang, or TensorRT-LLM
- Ability to read and modify serving engine internals
- Strong Python and PyTorch proficiency
- GPU profiling and optimization skills including CUDA and memory management
- Production ML systems shipped with measurable performance gains
- Product sense for developer needs
- Comfort on small, early-stage teams moving quickly
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or equivalent practical experience
Nice to have
- Speech and audio ML experience (ASR, TTS, signal processing)
- Familiarity with audio codecs and tokenization schemes
- Experience training or fine-tuning speech models