Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AISan Francisco, CA · full time · Staff
$250,000 – $300,000
Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Build and operate massive distributed storage systems for AI and machine learning workloads, architecting high-performance storage solutions that handle tens of petabytes of data and delivering extreme throughput to GPU clusters. This role suits engineers with deep expertise in storage systems and cloud infrastructure who want to solve infrastructure challenges at massive scale.
Our summary, not Together AI’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 8+ years in storage engineering managing multi-petabyte distributed storage
- Production experience deploying and operating high-performance storage for GPU or HPC clusters
- Deep Kubernetes and cloud-native storage experience in production
- Strong coding ability in Go and Python
- BS or MS in Computer Science, Engineering, or equivalent practical experience
- Demonstrated technical leadership improving system performance, reliability, or cost efficiency
- Deep expertise in at least one of: Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale
- Production experience with S3, MinIO, Ceph, or R2 object storage
- Experience with Kubernetes storage: CSI drivers, StatefulSets, PersistentVolumes, storage operators, custom controllers
- Storage optimization expertise for GPU workloads
- RDMA or InfiniBand networking knowledge
- Parallel filesystem optimization for TB/s+ aggregate cluster throughput
- Infrastructure as Code with Terraform, Ansible, Helm, or ArgoCD
- Advanced Linux storage knowledge: filesystems, LVM, NVMe, RAID
Nice to have
- GPU Direct Storage (GDS)
- NVMe-oF
- Storage networking expertise
- RDMA implementation experience
- ML and AI storage patterns like model weights, checkpointing, dataset caching
- Storage benchmarking and profiling tools expertise