$190,000 – $270,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
A systems engineering role focused on building and automating large-scale GPU infrastructure for AI model training and inference. This suits engineers who approach infrastructure as a software problem and are driven to eliminate manual operations through intelligent automation at massive scale.
Our summary, not Together AI’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 3+ years building distributed systems, infrastructure platforms, or large-scale backend software
- Strong software engineering skills in Python, Go, or Rust
- Experience building platforms, automation systems, or developer infrastructure
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies
- Systems thinking across hardware and software
- Passion for solving complex infrastructure challenges through software
- Automation-first mindset
Nice to have
- GPU infrastructure and CUDA experience
- NCCL, NVLink, or NVSwitch experience
- InfiniBand or RoCE networking experience
- Bare-metal provisioning and lifecycle management experience
- Large-scale AI training or inference cluster experience
- Hardware health monitoring and predictive failure detection experience
- Distributed storage systems experience
- AI agents and autonomous infrastructure operations experience