$152,000 – $287,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role optimizes and operates large-scale job scheduling systems (LSF, Slurm) that power EDA compute infrastructure across multiple sites, combining day-to-day operational management with improvements to observability, automation, and reliability. It suits experienced Linux systems engineers who enjoy solving complex infrastructure problems and working cross-functionally to translate technical decisions for non-technical stakeholders.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's degree in Computer Science or related field, or equivalent experience
- 5+ years operating large-scale Linux-based compute infrastructure
- Hands-on experience supporting and tuning job scheduling systems in HPC or silicon design environments
- Linux systems administration proficiency (CentOS/RHEL)
- Strong problem-solving skills and ability to independently analyze complex system behavior
- Clear communication of technical tradeoffs and reliability metrics to stakeholders
Nice to have
- Reliability engineering practices implementation in HPC scheduling environments
- Deep knowledge of scheduler configuration tuning, internals, and advanced troubleshooting
- Experience building or enhancing observability systems (metrics, monitoring, alerting, dashboards)
- Container technologies experience (Docker, Singularity, Podman) in HPC
- Experience influencing infrastructure standards adoption across multiple teams or sites