# Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Hiring organization: [Together AI](https://career.thegoodapps.co/organizations/together-ai)

Canonical page: https://career.thegoodapps.co/jobs/bbf5575a-969d-42b0-8981-d8a60f08236c

Listed on Together AI's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Staff
- Location: San Francisco, CA
- Salary: 250000 – 300000 USD per year

## Summary

Build and operate massive distributed storage systems for AI and machine learning workloads, architecting high-performance storage solutions that handle tens of petabytes of data and delivering extreme throughput to GPU clusters. This role suits engineers with deep expertise in storage systems and cloud infrastructure who want to solve infrastructure challenges at massive scale.

_Our summary, not Together AI's wording._

## Skills named

Amazon S3, Ansible, ArgoCD, Go, Grafana, Helm, Kubernetes, Prometheus, Python, Terraform, Thanos, Weka

## Required

- 8+ years in storage engineering managing multi-petabyte distributed storage
- Production experience deploying and operating high-performance storage for GPU or HPC clusters
- Deep Kubernetes and cloud-native storage experience in production
- Strong coding ability in Go and Python
- BS or MS in Computer Science, Engineering, or equivalent practical experience
- Demonstrated technical leadership improving system performance, reliability, or cost efficiency
- Deep expertise in at least one of: Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale
- Production experience with S3, MinIO, Ceph, or R2 object storage
- Experience with Kubernetes storage: CSI drivers, StatefulSets, PersistentVolumes, storage operators, custom controllers
- Storage optimization expertise for GPU workloads
- RDMA or InfiniBand networking knowledge
- Parallel filesystem optimization for TB/s+ aggregate cluster throughput
- Infrastructure as Code with Terraform, Ansible, Helm, or ArgoCD
- Advanced Linux storage knowledge: filesystems, LVM, NVMe, RAID

## Nice to have

- GPU Direct Storage (GDS)
- NVMe-oF
- Storage networking expertise
- RDMA implementation experience
- ML and AI storage patterns like model weights, checkpointing, dataset caching
- Storage benchmarking and profiling tools expertise

Apply on Together AI's site: https://job-boards.greenhouse.io/togetherai/jobs/5155722007
