$224,000 – $431,250
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
NVIDIA is seeking a senior infrastructure engineer to build and optimize distributed training systems for autonomous vehicle AI models running across massive GPU clusters. This role suits experienced systems engineers who excel at scaling complex distributed systems and want to work on foundational ML infrastructure at one of the world's leading AI companies.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or related field, or equivalent experience
- 12+ years building and scaling high-performance distributed systems in ML, HPC, or large-scale data infrastructure
- Deep expertise with PyTorch and large-scale training techniques
- Strong systems knowledge including datacenter networking, parallel filesystems, and schedulers
- Production-grade Python library development experience
- Ability to collaborate effectively with ML researchers and infrastructure teams
Nice to have
- Experience scaling GPU training clusters with 1,000+ GPUs
- Expertise in fault resilience, high availability, and elastic training
- Large-scale observability and monitoring systems
- Technical leadership experience as a hands-on authority in ML systems engineering