# Research Engineer, Large-Scale Training

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/ea81e48c-0c1d-4280-859f-b31394dcdb75

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Location: San Francisco, CA
- Salary: 200000 – 290000 USD per year

## Summary

A Research Engineer role focused on building and optimizing large-scale training infrastructure for foundation models at Together AI. This suits engineers who combine systems expertise with ML knowledge and enjoy translating research into production systems that serve real customers.

_Our summary, not Together AI's wording._

## Skills named

CUDA, Python, PyTorch

## Required

- Ability to independently investigate and deploy solutions to performance and infrastructure problems
- Strong Python and PyTorch programming with focus on efficiency and maintainability
- Hands-on experience training or fine-tuning large neural networks on multi-GPU or multi-node setups
- Understanding of ML systems fundamentals including GPU architecture, mixed-precision training, and distributed training paradigms
- Strong communication and collaboration skills with researchers and engineers
- Commitment to staying current with AI research advances

## Nice to have

- Experience writing optimized NVIDIA GPU kernels using CUDA or Triton
- Experience with NCCL or NVSHMEM communication collectives
- Experience with large-scale training frameworks like FSDP, DeepSpeed, or Megatron-LM
- Experience optimizing distributed training for compute, memory, or scalability efficiency
- Experience running and managing large-scale GPU experiments with scheduling and fault tolerance
- Open-source ML or ML systems contributions
- Experience building or operating ML products or managed services for external customers

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/5199554007
