# Senior System Software Engineer - GPU Performance

Hiring organization: [NVIDIA](https://career.thegoodapps.co/organizations/nvidia)

Canonical page: https://career.thegoodapps.co/jobs/be81e7d9-d170-41af-ac57-d50e4e726a2d

Listed on NVIDIA's own careers site. Applications go to them directly.

- Seniority: Senior
- Salary: 152000 – 287500 USD per year

## Summary

NVIDIA seeks a performance engineer to optimize communication libraries (NCCL, NVSHMEM, UCX) that connect thousands of GPUs in deep learning and HPC systems. This role suits someone with systems software expertise who enjoys diagnosing performance bottlenecks across GPU clusters and networking stacks.

_Our summary, not NVIDIA's wording._

## Skills named

Ansible, Automotive Ethernet, CUDA, Docker, Kubernetes, Python, PyTorch, TensorFlow

## Required

- Master's degree or PhD in Computer Science or related field
- 3+ years parallel programming experience
- Experience with at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
- Performance benchmarking and triage on large-scale HPC clusters
- Understanding of computer system architecture and HW-SW interactions
- Implement micro-benchmarks in C/C++
- Debug performance issues across HW/SW stack
- Proficiency in a scripting language, preferably Python
- Familiarity with containers, cloud provisioning and scheduling tools

## Nice to have

- Practical experience with Infiniband/Ethernet networks, RDMA, topologies, and congestion control
- Experience debugging network issues in large-scale deployments
- CUDA programming and/or GPU familiarity
- Experience with Deep Learning Frameworks such as PyTorch or TensorFlow

Apply on NVIDIA's site: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Senior-HPC-Performance-Engineer_JR1997214
