$200,000 – $290,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
A Research Engineer role focused on building and optimizing large-scale training infrastructure for foundation models at Together AI. This suits engineers who combine systems expertise with ML knowledge and enjoy translating research into production systems that serve real customers.
Our summary, not Together AI’s wording. The full posting is on their site.
What they ask for
Required
- Ability to independently investigate and deploy solutions to performance and infrastructure problems
- Strong Python and PyTorch programming with focus on efficiency and maintainability
- Hands-on experience training or fine-tuning large neural networks on multi-GPU or multi-node setups
- Understanding of ML systems fundamentals including GPU architecture, mixed-precision training, and distributed training paradigms
- Strong communication and collaboration skills with researchers and engineers
- Commitment to staying current with AI research advances
Nice to have
- Experience writing optimized NVIDIA GPU kernels using CUDA or Triton
- Experience with NCCL or NVSHMEM communication collectives
- Experience with large-scale training frameworks like FSDP, DeepSpeed, or Megatron-LM
- Experience optimizing distributed training for compute, memory, or scalability efficiency
- Experience running and managing large-scale GPU experiments with scheduling and fault tolerance
- Open-source ML or ML systems contributions
- Experience building or operating ML products or managed services for external customers