# Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

Hiring organization: [NVIDIA](https://career.thegoodapps.co/organizations/nvidia)

Canonical page: https://career.thegoodapps.co/jobs/4e95e136-0ef2-4904-84f6-8fcb70c34fbe

Listed on NVIDIA's own careers site. Applications go to them directly.

- Seniority: Principal
- Location: Santa Clara, CA
- Salary: 272000 – 431250 USD per year

## Summary

A principal-level engineer focused on making NVIDIA's systems reliable at hyperscale, working with major cloud providers to measure and improve mean-time-between-interruptions for GPU platforms in production. This role suits someone with deep datacenter systems expertise and a track record of translating fleet reliability data into actionable improvements across firmware, drivers, and hardware.

_Our summary, not NVIDIA's wording._

## Skills named

Anomaly Detection, Medical Telemetry, Root Cause Analysis, Statistical Analysis, Thermal Management

## Required

- 15+ years systems software or reliability engineering at datacenter scale
- BS or MS in Computer Science, Electrical Engineering, Statistics, or equivalent experience
- Deep expertise in multi-NUMA and rack-scale system software and firmware
- Statistical failure analysis methods including MTBF/MTBI calculation and Pareto analysis
- Fleet-level telemetry and observability systems experience
- Understanding of hardware failure modes in large-scale GPU/accelerator deployments
- Experience defining or operating burn-in, stress testing, or certification frameworks
- Customer obsession and genuine passion for fleet reliability challenges at scale
- Strong communication skills for presenting findings to technical and executive audiences
- Demonstrated success driving cross-functional improvements without direct authority

## Nice to have

- Fleet reliability experience at a hyperscaler
- Familiarity with NVIDIA GPU error taxonomy
- Experience building health scoring or predictive failure models for accelerator or HPC infrastructure
- Background defining MTBI/MTBF measurement standards or certification programs
- Understanding of reliability data flow from device firmware through telemetry pipelines to fleet dashboards

Apply on NVIDIA's site: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--At-Scale-Reliability-and-Fleet-Intelligence---CSP-Engagements_JR2020320
