Skip to main content
CareerApp

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AI

San Francisco, CA · full time · Staff

$250,000 – $300,000

Listed on Together AI’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Build and operate massive distributed storage systems for AI and machine learning workloads, architecting high-performance storage solutions that handle tens of petabytes of data and delivering extreme throughput to GPU clusters. This role suits engineers with deep expertise in storage systems and cloud infrastructure who want to solve infrastructure challenges at massive scale.

Our summary, not Together AI’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 8+ years in storage engineering managing multi-petabyte distributed storage
  • Production experience deploying and operating high-performance storage for GPU or HPC clusters
  • Deep Kubernetes and cloud-native storage experience in production
  • Strong coding ability in Go and Python
  • BS or MS in Computer Science, Engineering, or equivalent practical experience
  • Demonstrated technical leadership improving system performance, reliability, or cost efficiency
  • Deep expertise in at least one of: Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale
  • Production experience with S3, MinIO, Ceph, or R2 object storage
  • Experience with Kubernetes storage: CSI drivers, StatefulSets, PersistentVolumes, storage operators, custom controllers
  • Storage optimization expertise for GPU workloads
  • RDMA or InfiniBand networking knowledge
  • Parallel filesystem optimization for TB/s+ aggregate cluster throughput
  • Infrastructure as Code with Terraform, Ansible, Helm, or ArgoCD
  • Advanced Linux storage knowledge: filesystems, LVM, NVMe, RAID

Nice to have

  • GPU Direct Storage (GDS)
  • NVMe-oF
  • Storage networking expertise
  • RDMA implementation experience
  • ML and AI storage patterns like model weights, checkpointing, dataset caching
  • Storage benchmarking and profiling tools expertise

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.