$272,000 – $431,250
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
NVIDIA seeks a Principal engineer to architect and oversee the software systems that manage rack-scale infrastructure products, working across firmware, operating systems, control planes, and networking to convert complex hardware into reliable services for internal teams and cloud customers worldwide. This role demands deep expertise in distributed systems and infrastructure software, with the opportunity to mentor senior engineers and shape how NVIDIA's infrastructure scales across the data center.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 15+ years in systems architecture, system software, distributed systems, infrastructure control planes, or infrastructure engineering
- BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or equivalent experience
- Architectural knowledge of coordination frameworks, state machines, declarative APIs, reconciliation loops, and distributed systems tradeoffs
- Production-quality coding in Go, C++, or Rust
- Experience with Kubernetes or similar orchestration systems for managing infrastructure at scale
- Data center networking knowledge including Ethernet, InfiniBand, RDMA, and fabric-level manageability
- Experience with accelerator-based systems like GPUs, DPUs, FPGAs, or custom silicon
- Understanding of in-band and out-of-band management including BMCs, Redfish, and IPMI
- Experience architecting software for open source release with API stability and modularity
- Cross-team technical leadership and clear communication of hardware/software tradeoffs
Nice to have
- Experience with Rust in systems or infrastructure software
- Built software supporting multiple adoption models across internal services, CSP offerings, libraries, and customer APIs
- Fleet-scale provisioning, updates, rollback, observability, and remediation at scale
- Led products through full lifecycle from inception through post-silicon, manufacturing, deployment, and operations
- Experience with open source ecosystems and balancing community collaboration with product needs
- Deep rack- or cluster-scale systems experience spanning compute, networking, storage, accelerators, and firmware as one domain
- Skilled at creating simple abstractions across complex systems
- Experience using AI-assisted development tools responsibly