Skip to main content
CareerApp
NVIDIA

Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

NVIDIA

Santa Clara, CA · Principal

$272,000 – $431,250

Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

A principal-level engineer focused on making NVIDIA's systems reliable at hyperscale, working with major cloud providers to measure and improve mean-time-between-interruptions for GPU platforms in production. This role suits someone with deep datacenter systems expertise and a track record of translating fleet reliability data into actionable improvements across firmware, drivers, and hardware.

Our summary, not NVIDIA’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 15+ years systems software or reliability engineering at datacenter scale
  • BS or MS in Computer Science, Electrical Engineering, Statistics, or equivalent experience
  • Deep expertise in multi-NUMA and rack-scale system software and firmware
  • Statistical failure analysis methods including MTBF/MTBI calculation and Pareto analysis
  • Fleet-level telemetry and observability systems experience
  • Understanding of hardware failure modes in large-scale GPU/accelerator deployments
  • Experience defining or operating burn-in, stress testing, or certification frameworks
  • Customer obsession and genuine passion for fleet reliability challenges at scale
  • Strong communication skills for presenting findings to technical and executive audiences
  • Demonstrated success driving cross-functional improvements without direct authority

Nice to have

  • Fleet reliability experience at a hyperscaler
  • Familiarity with NVIDIA GPU error taxonomy
  • Experience building health scoring or predictive failure models for accelerator or HPC infrastructure
  • Background defining MTBI/MTBF measurement standards or certification programs
  • Understanding of reliability data flow from device firmware through telemetry pipelines to fleet dashboards

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.