$184,000 – $356,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves integrating CUDA features and distributed runtime technologies into AI frameworks like PyTorch and vLLM, working on the performance and scalability of deep learning systems across multi-GPU clusters. It suits experienced systems engineers who combine deep learning knowledge with low-level optimization expertise and want to shape the infrastructure that powers modern AI.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's, Master's or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field
- 8+ years of industry experience or equivalent after degree
- Development experience with deep learning frameworks like PyTorch and JAX
- Development experience with inference engines like TRT-LLM, vLLM, or sgLang
- Rapid prototyping in Python, C++, or CUDA
- Understanding of AI models, parallelism, and compiler technologies
- Performance benchmarking experience on AI clusters
- Familiarity with PyTorch profiler or NVIDIA Nsight Systems
- Understanding of HPC/AI communication concepts
- Knowledge of computer architecture, hardware-software interactions, and operating systems
Nice to have
- Deep expertise in performance internals of PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron, or MaxText
- Hands-on experience with NCCL, MPI, UCX, and distributed ML techniques like pipeline or tensor parallelism
- Expertise in training, distributed inference, MoE, reinforcement learning, or kernel authoring
- Background in deep learning compilers at graph or codegen level
- Experience with compute and communication overlap in distributed systems