# Senior Deep Learning Software Infrastructure Engineer

Hiring organization: [NVIDIA](https://career.thegoodapps.co/organizations/nvidia)

Canonical page: https://career.thegoodapps.co/jobs/b2dffd1c-f60f-495b-90dd-6b16f95b33e1

Listed on NVIDIA's own careers site. Applications go to them directly.

- Seniority: Senior
- Location: US, CA, Remote
- Remote: yes
- Salary: 224000 – 431250 USD per year

## Summary

NVIDIA is seeking a senior infrastructure engineer to build and optimize distributed training systems for autonomous vehicle AI models running across massive GPU clusters. This role suits experienced systems engineers who excel at scaling complex distributed systems and want to work on foundational ML infrastructure at one of the world's leading AI companies.

_Our summary, not NVIDIA's wording._

## Skills named

CUDA, Kubernetes, Python, PyTorch

## Required

- Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, Computer Engineering, or related field, or equivalent experience
- 12+ years building and scaling high-performance distributed systems in ML, HPC, or large-scale data infrastructure
- Deep expertise with PyTorch and large-scale training techniques
- Strong systems knowledge including datacenter networking, parallel filesystems, and schedulers
- Production-grade Python library development experience
- Ability to collaborate effectively with ML researchers and infrastructure teams

## Nice to have

- Experience scaling GPU training clusters with 1,000+ GPUs
- Expertise in fault resilience, high availability, and elastic training
- Large-scale observability and monitoring systems
- Technical leadership experience as a hands-on authority in ML systems engineering

Apply on NVIDIA's site: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-Deep-Learning-Sofware-Infrastructure-Engineer_JR2022034
