$184,000 – $356,500
Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
NVIDIA is hiring a senior engineer to provide hands-on technical support for large-scale AI deployments on their cloud platforms, working directly with enterprise customers to optimize distributed training, inference, and MLOps infrastructure. This role suits experienced infrastructure engineers or solutions architects comfortable troubleshooting complex production systems and guiding customer teams on GPU-accelerated platforms.
Our summary, not NVIDIA’s wording. The full posting is on their site.
Skills this role names
- Ansible
- CI/CD
- CUDA
- Distributed Computing
- Docker
- Go
- Grafana
- Kubernetes
- Linux
- MLOps
- Nim
- OpenTelemetry
- Prometheus
- Python
- PyTorch
- TensorFlow
- Terraform
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or related technical field, or equivalent experience
- 8+ years in customer-facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering
- Expertise in Linux systems, distributed computing, Kubernetes, containers, and GPU scheduling on multi-tenant platforms
- Experience supporting large-scale AI/ML training and inference workloads in production environments
- Programming proficiency in Python or Go with hands-on framework experience (PyTorch or TensorFlow)
- Ability to collaborate with customer and partner engineering teams, lead technical investigations, and resolve issues
- Strong communication and technical presentation skills
Nice to have
- Experience with NVIDIA ecosystem tools (DGX, CUDA, NeMo, Triton, NIM, InfiniBand, RoCE)
- Direct collaboration experience with NVIDIA Cloud Partners, hyperscalers, or managed AI platforms
- Deep familiarity with MLOps, cloud-native practices, containerization, CI/CD pipelines, and observability tools
- Background in infrastructure as code (Terraform, Ansible, or similar) for GPU cluster deployment