Staff Software Engineer, Inference / Compute Infrastructure Engineering
Together AISan Francisco, CA · full time · Lead / Staff / Principal
$240,000 – $280,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Design and build the Kubernetes control plane that manages Together AI's GPU inference fleet, creating declarative APIs so the inference team can provision and scale clusters without wrestling with the underlying infrastructure complexity. You'll own the full stack from provisioning logic through scheduling optimization and self-healing systems, combining infrastructure engineering with product thinking to make the platform seamless for internal customers.
Our summary, not Together AI’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong software engineering background in Go, Python, Rust or equivalent
- Experience with durable workflow orchestration tools like Temporal or Cadence
- Experience building software control planes or orchestration systems that model and reconcile state
- Experience designing and building event-driven systems with message queues or pub/sub
- Product mindset with experience building internal platforms consumed by other engineering teams
Nice to have
- Bare-metal provisioning experience (PXE/iPXE, Redfish/IPMI, BMC)
- Networking fundamentals (VLANs, BGP, fabric design)
- GPU or accelerator infrastructure experience
- GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE)
- Prior experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization
- Systems programming in Rust or Go