Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves designing and operating the Kubernetes platform that underpins NVIDIA's global network infrastructure, with responsibility for cluster lifecycle, automation, and production support. It suits experienced platform engineers who want deep ownership of mission-critical systems and can bridge complex distributed systems work with network operations at scale.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's degree in Computer Science, Engineering, or related field or equivalent experience
- 8+ years building or operating production Kubernetes platforms, network infrastructure, or distributed systems
- Deep Kubernetes expertise including cluster lifecycle, upgrades, networking, storage, and recovery
- Proficiency in Go or Python
- Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery
- Experience deploying and supporting network automation or telemetry services on Kubernetes
- Production on-call, incident response, root-cause analysis, and corrective action experience
Nice to have
- Knowledge of IP routing, data center fabrics, and cloud networking
- Experience designing and operating large multi-region Kubernetes fleets with fleet-wide upgrades and recovery
- Hands-on experience with Cluster API and Metal3 for bare-metal provisioning and cluster lifecycle
- Experience building Kubernetes controllers or operators in Go
- Experience designing or operating network automation and telemetry services at global scale
- Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects