# AI Infrastructure Systems Engineer

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/b67d79be-b7d3-4de8-a1e3-51bdbd4a8364

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Mid
- Location: San Francisco, CA
- Salary: 190000 – 270000 USD per year

## Summary

A systems engineering role focused on building and automating large-scale GPU infrastructure for AI model training and inference. This suits engineers who approach infrastructure as a software problem and are driven to eliminate manual operations through intelligent automation at massive scale.

_Our summary, not Together AI's wording._

## Skills named

Ansible, CUDA, Go, Kubernetes, Linux, Python, Rust, Terraform

## Required

- 3+ years building distributed systems, infrastructure platforms, or large-scale backend software
- Strong software engineering skills in Python, Go, or Rust
- Experience building platforms, automation systems, or developer infrastructure
- Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies
- Systems thinking across hardware and software
- Passion for solving complex infrastructure challenges through software
- Automation-first mindset

## Nice to have

- GPU infrastructure and CUDA experience
- NCCL, NVLink, or NVSwitch experience
- InfiniBand or RoCE networking experience
- Bare-metal provisioning and lifecycle management experience
- Large-scale AI training or inference cluster experience
- Hardware health monitoring and predictive failure detection experience
- Distributed storage systems experience
- AI agents and autonomous infrastructure operations experience

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/5138540007
