$152,000 – $287,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
NVIDIA seeks a deep learning engineer to integrate advanced communication technologies into AI frameworks like PyTorch and JAX, optimizing multi-GPU performance for training and inference workloads. The role involves hands-on development with communication libraries, compiler improvements, and performance analysis across large-scale distributed systems.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's degree or higher in Computer Science or related field, or equivalent experience
- 5+ years of software engineering and HPC/AI experience
- Development or integration with deep learning frameworks like PyTorch or JAX
- Experience with inference engines such as TRT-LLM, vLLM, or SGLang
- Proficiency in Python, C++, CUDA, or related domain-specific languages
- Understanding of AI models, parallelism, and compiler technologies
- Performance benchmarking experience on AI clusters with profiling tools
- Knowledge of HPC/AI communication concepts including 1-sided/2-sided communication, elasticity, and resiliency
Nice to have
- Experience with parallel programming runtimes like NCCL, NVSHMEM, or MPI
- Strong understanding of computer system architecture and OS principles
- Expertise in training, distributed inference, MoE, or reinforcement learning
- Kernel authoring experience with CUDA, Triton, or cuTe
- Experience with compute-communication overlap in distributed runtimes
- AI compiler pattern matching and lowering expertise
- Understanding of memory hierarchy, consistency models, and tensor layouts