$209,000 – $253,000
Listed on Crusoe’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Crusoe seeks a senior production engineer to lead reliability and operational excellence across their GPU cloud infrastructure that powers AI workloads. This role combines technical depth in distributed systems and infrastructure with leadership responsibilities, mentoring junior engineers while architecting observability and automation solutions for large-scale, latency-sensitive environments.
Our summary, not Crusoe’s wording. The full posting is on their site.
Skills this role names
- Amazon Web Services (AWS)
- Ansible
- C++
- Go
- Grafana
- Kubernetes
- Linux
- OpenTelemetry
- Prometheus
- Python
- Terraform
- Unix
Log in to see which of these are already on your profile.
What they ask for
Required
- Bachelor's degree in Computer Science, Engineering, or related technical field or equivalent practical experience
- 8+ years of Production Engineering, SRE, or large-scale infrastructure operations experience
- Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems
- Previous experience in infrastructure roles building or managing compute, storage or networking platforms
- Deep knowledge of Linux/Unix systems including kernel and user space debugging
- Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms
- Track record with incident management practices and reliability frameworks
- Hands-on experience with Prometheus and Grafana
- Experience with infrastructure-as-code and configuration management tools
- Proficiency in scripting or programming with Go, Python, C, or C++
- Exceptional communication skills and cross-team collaboration ability
- Ability to remain effective while troubleshooting complex production issues
- Commitment to reliability engineering and operational excellence
Nice to have
- Experience leading Kubernetes or container orchestration platforms at scale
- Exposure to change management processes, operational readiness reviews, or structured root cause analysis
- Experience designing self-healing systems, automated remediation, or event-driven operational tooling
- Interest in scaling AI or HPC infrastructure and GPU-heavy reliability challenges
- Passion for mentorship and developing Production Engineering expertise