# Automated Testing Engineer, Compute

Hiring organization: [Crusoe](https://career.thegoodapps.co/organizations/crusoe)

Canonical page: https://career.thegoodapps.co/jobs/d10176fe-952d-4ec0-a218-c62f744ebf1c

Listed on Crusoe's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Mid
- Location: San Francisco, CA
- Salary: 172500 – 210000 USD per year

## Summary

This role involves designing and maintaining automated testing frameworks for large-scale GPU clusters, with a focus on validating interconnect performance, distributed workload scaling, and multi-node stability in virtualized environments. It suits engineers with deep systems knowledge who enjoy working on low-level infrastructure challenges in a fast-moving AI compute company.

_Our summary, not Crusoe's wording._

## Skills named

Bash (Scripting), CUDA, Docker, GitLab, Kubernetes, Linux Kernel, Python, Terraform

## Required

- 5+ years of experience in relevant role
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
- Experience building and deploying automated integration testing for AI cloud environments
- Working knowledge of Kubernetes, Docker, Terraform, and Postgres
- Advanced proficiency in Python and/or Bash
- Familiarity with NVIDIA CUDA/NCCL and/or AMD ROCm/RCCL stacks in multi-node context
- Strong understanding of RDMA, RoCE, and InfiniBand in virtualized systems
- Knowledge of Linux kernel internals, PCIe topology, VFIO, and memory management

## Nice to have

- Experience with MNNVL or specialized AI fabric architectures
- Familiarity with hardware-level debugging tools and performance profilers like NVIDIA Nsight or AMD Omniperf
- Knowledge of containerized GPU orchestration with Kubernetes device plugins

Apply on Crusoe's site: https://jobs.ashbyhq.com/Crusoe/beb97388-6fff-427d-97c6-ae7ea1385d58/application
