$184,000 – $287,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
A research-focused software engineering role optimizing deep learning systems at the intersection of CUDA and AI workloads, suitable for engineers with deep systems expertise who want to tackle hardware-software co-design challenges across single GPUs to supercomputer clusters.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field
- 8+ years of relevant industry experience or equivalent academic experience
- Strong proficiency in C++ and Python
- Deep learning fundamentals with focus on transformers
- Distributed computing and multi-node scaling knowledge
- Systems programming and computer architecture background
- Low-level systems performance optimization experience
- GPU and CUDA programming with kernel optimization experience
- Experience profiling and optimizing generative AI models
Nice to have
- Deep expertise in PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron internals
- Hands-on experience with NCCL, MPI, UCX communication libraries
- Knowledge of low-precision arithmetic (NVFP4, MXFP4, FP8, INT8)
- Background in deep learning compilers and ML systems (Triton, XLA, torch.compile)
- Experience designing agentic AI systems for infrastructure problems