$172,500 – $210,000
Listed on Crusoe’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves designing and maintaining automated testing frameworks for large-scale GPU clusters, with a focus on validating interconnect performance, distributed workload scaling, and multi-node stability in virtualized environments. It suits engineers with deep systems knowledge who enjoy working on low-level infrastructure challenges in a fast-moving AI compute company.
Our summary, not Crusoe’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 5+ years of experience in relevant role
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
- Experience building and deploying automated integration testing for AI cloud environments
- Working knowledge of Kubernetes, Docker, Terraform, and Postgres
- Advanced proficiency in Python and/or Bash
- Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL stacks in multi-node context
- Strong understanding of RDMA, RoCE, and InfiniBand in virtualized systems
- Knowledge of Linux kernel internals, PCIe topology, VFIO, and memory management
Nice to have
- Experience with MNNVL or specialized AI fabric architectures
- Familiarity with hardware-level debugging tools and performance profilers like NVIDIA Nsight or AMD Omniperf
- Knowledge of containerized GPU orchestration with Kubernetes device plugins