$184,000 – $287,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
NVIDIA is seeking a senior engineer to architect and debug cloud-native software stacks for AI datacenters running GB200 and GB300 GPUs, working directly with cloud providers to solve large-scale Kubernetes and Slurm challenges. This role suits someone with deep distributed-systems expertise who enjoys both hands-on debugging and customer-facing technical leadership.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 10+ years professional software development in distributed systems
- Source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators)
- Source-level expertise in Slurm (federation, power-save, plugins)
- Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters
- Proven track record debugging large-scale cloud-native stacks across networking, storage, and control planes
- Customer-facing engineering or solutions-architect background with requirements gathering and PoC ownership
- BS or MS in Computer Engineering, Computer Science, or related field
Nice to have
- Upstream contributions to Kubernetes, Slurm, Volcano, or similar projects
- Experience with GPU computing (CUDA) and deep learning workloads
- Familiarity with CI/CD, observability tools, and infrastructure-as-code