$152,000 – $287,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
NVIDIA seeks a performance engineer to optimize communication libraries (NCCL, NVSHMEM, UCX) that connect thousands of GPUs in deep learning and HPC systems. This role suits someone with systems software expertise who enjoys diagnosing performance bottlenecks across GPU clusters and networking stacks.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Master's degree or PhD in Computer Science or related field
- 3+ years parallel programming experience
- Experience with at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
- Performance benchmarking and triage on large-scale HPC clusters
- Understanding of computer system architecture and HW-SW interactions
- Implement micro-benchmarks in C/C++
- Debug performance issues across HW/SW stack
- Proficiency in a scripting language, preferably Python
- Familiarity with containers, cloud provisioning and scheduling tools
Nice to have
- Practical experience with Infiniband/Ethernet networks, RDMA, topologies, and congestion control
- Experience debugging network issues in large-scale deployments
- CUDA programming and/or GPU familiarity
- Experience with Deep Learning Frameworks such as PyTorch or TensorFlow