# Staff Software Engineer, Inference / Compute Infrastructure Engineering

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/038d19af-6484-416b-a7cb-11544a1401f9

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Lead / Staff / Principal
- Location: San Francisco, CA
- Salary: 240000 – 280000 USD per year

## Summary

Design and build the Kubernetes control plane that manages Together AI's GPU inference fleet, creating declarative APIs so the inference team can provision and scale clusters without wrestling with the underlying infrastructure complexity. You'll own the full stack from provisioning logic through scheduling optimization and self-healing systems, combining infrastructure engineering with product thinking to make the platform seamless for internal customers.

_Our summary, not Together AI's wording._

## Skills named

Apache Kafka, Border Gateway Protocol (BGP), Cadence OrCAD/Allegro, CUDA, Go, Kubernetes, Python, Rust

## Required

- Strong software engineering background in Go, Python, Rust or equivalent
- Experience with durable workflow orchestration tools like Temporal or Cadence
- Experience building software control planes or orchestration systems that model and reconcile state
- Experience designing and building event-driven systems with message queues or pub/sub
- Product mindset with experience building internal platforms consumed by other engineering teams

## Nice to have

- Bare-metal provisioning experience (PXE/iPXE, Redfish/IPMI, BMC)
- Networking fundamentals (VLANs, BGP, fabric design)
- GPU or accelerator infrastructure experience
- GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE)
- Prior experience at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization
- Systems programming in Rust or Go

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/5186628007
