Principal Software Engineer, Rack-Scale System Software — CSP Engagements
NVIDIAfull time · Principal
$272,000 – $431,250
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role leads technical engagement with cloud service providers on NVIDIA's rack-scale system software and firmware, serving as the bridge between CSP operations teams and NVIDIA's internal engineering. You'll drive architecture decisions, synthesize customer feedback into product improvements, and ensure complex multi-component systems are reliable and operable at fleet scale.
Our summary, not NVIDIA’s wording. The full posting is on their site.
What they ask for
Required
- 15+ years in system software, platform firmware, or large-scale distributed systems
- BS/MS in Computer Science, Electrical Engineering, or equivalent experience
- Deep understanding of rack-scale system software challenges including multi-component coordination and error propagation
- Experience with fabric management software, cluster management, or system-level orchestration frameworks
- Understanding of firmware architectures and update lifecycle management
- Knowledge of error handling and recovery design patterns in distributed systems
- Experience with health monitoring and telemetry systems
- Technical leadership across organizational boundaries without direct authority
- Strong communication skills translating complex architecture to customer engineering teams
Nice to have
- GPU or accelerator system software experience (drivers, device management, power management)
- Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software
- Background in system software for large-scale clusters at a hyperscaler
- Experience crafting error handling and recovery frameworks for multi-component systems
- Familiarity with GPU or accelerator fleet operations including driver lifecycle and firmware rollout strategies
- Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale