# Senior Solutions Architect, AI Factory Deployment - NVIS

Hiring organization: [NVIDIA](https://career.thegoodapps.co/organizations/nvidia)

Canonical page: https://career.thegoodapps.co/jobs/a38885d7-3d5d-4525-bf46-877d6b4f8b52

Listed on NVIDIA's own careers site. Applications go to them directly.

- Seniority: Senior
- Salary: 152000 – 287500 USD per year

## Summary

This role involves setting up, validating, and troubleshooting AI factories running on GPU clusters, with a focus on optimizing distributed training workloads through NCCL collectives and performance analysis. It suits experienced systems engineers or performance specialists who combine deep Linux and HPC knowledge with Python automation skills and want to work on the infrastructure layer that powers large-scale AI training.

_Our summary, not NVIDIA's wording._

## Skills named

Bash (Scripting), Linux, Python

## Required

- Bachelor's degree or equivalent in Computer Science, Mathematics, Engineering, Physics, or related field
- 5+ years managing Linux-based systems in HPC, distributed systems, or AI/ML environments
- Hands-on experience running AI/ML workloads on multi-GPU and/or multi-node clusters with some NCCL exposure
- Practical knowledge of collective communication patterns like AllReduce and AllToAll in ML/LLM training
- Python and Shell/Bash proficiency for scripting and automation
- Strong communication and cross-functional collaboration skills

## Nice to have

- Experience benchmarking distributed systems
- Background in HPC performance engineering, SRE, or GPU-accelerated systems performance analysis
- Familiarity with observability stacks for large distributed systems
- Experience building automation and CI-style pipelines for benchmark validation at scale
- Demonstrated interest in using AI to solve practical problems and guide data-driven decisions

Apply on NVIDIA's site: https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCareerSite/job/US-CA-Remote/Senior-Solutions-Architect--AI-Factory-Deployment---NVIS_JR2022471
