# Principal Software Engineer, Rack-Scale System Software — CSP Engagements

Hiring organization: [NVIDIA](https://career.thegoodapps.co/organizations/nvidia)

Canonical page: https://career.thegoodapps.co/jobs/6bb16171-7c16-4b13-897a-c8e71e6f84cd

Listed on NVIDIA's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Principal
- Salary: 272000 – 431250 USD per year

## Summary

This role leads technical engagement with cloud service providers on NVIDIA's rack-scale system software and firmware, serving as the bridge between CSP operations teams and NVIDIA's internal engineering. You'll drive architecture decisions, synthesize customer feedback into product improvements, and ensure complex multi-component systems are reliable and operable at fleet scale.

_Our summary, not NVIDIA's wording._

## Required

- 15+ years in system software, platform firmware, or large-scale distributed systems
- BS/MS in Computer Science, Electrical Engineering, or equivalent experience
- Deep understanding of rack-scale system software challenges including multi-component coordination and error propagation
- Experience with fabric management software, cluster management, or system-level orchestration frameworks
- Understanding of firmware architectures and update lifecycle management
- Knowledge of error handling and recovery design patterns in distributed systems
- Experience with health monitoring and telemetry systems
- Technical leadership across organizational boundaries without direct authority
- Strong communication skills translating complex architecture to customer engineering teams

## Nice to have

- GPU or accelerator system software experience (drivers, device management, power management)
- Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software
- Background in system software for large-scale clusters at a hyperscaler
- Experience crafting error handling and recovery frameworks for multi-component systems
- Familiarity with GPU or accelerator fleet operations including driver lifecycle and firmware rollout strategies
- Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale

Apply on NVIDIA's site: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Principal-Software-Engineer--Rack-Scale-System-Software---CSP-Engagements_JR2020316
