# Staff Production Engineer, Core PE

Hiring organization: [Crusoe](https://career.thegoodapps.co/organizations/crusoe)

Canonical page: https://career.thegoodapps.co/jobs/309ea983-1b71-4b2a-9493-8009caa00060

Listed on Crusoe's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Staff
- Location: Sunnyvale, CA
- Salary: 209000 – 253000 USD per year

## Summary

Crusoe seeks a senior production engineer to lead reliability and operational excellence across their GPU cloud infrastructure that powers AI workloads. This role combines technical depth in distributed systems and infrastructure with leadership responsibilities, mentoring junior engineers while architecting observability and automation solutions for large-scale, latency-sensitive environments.

_Our summary, not Crusoe's wording._

## Skills named

Amazon Web Services (AWS), Ansible, C++, Go, Grafana, Kubernetes, Linux, OpenTelemetry, Prometheus, Python, Terraform, Unix

## Required

- Bachelor's degree in Computer Science, Engineering, or related technical field or equivalent practical experience
- 8+ years of Production Engineering, SRE, or large-scale infrastructure operations experience
- Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems
- Previous experience in infrastructure roles building or managing compute, storage or networking platforms
- Deep knowledge of Linux/Unix systems including kernel and user space debugging
- Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms
- Track record with incident management practices and reliability frameworks
- Hands-on experience with Prometheus and Grafana
- Experience with infrastructure-as-code and configuration management tools
- Proficiency in scripting or programming with Go, Python, C, or C++
- Exceptional communication skills and cross-team collaboration ability
- Ability to remain effective while troubleshooting complex production issues
- Commitment to reliability engineering and operational excellence

## Nice to have

- Experience leading Kubernetes or container orchestration platforms at scale
- Exposure to change management processes, operational readiness reviews, or structured root cause analysis
- Experience designing self-healing systems, automated remediation, or event-driven operational tooling
- Interest in scaling AI or HPC infrastructure and GPU-heavy reliability challenges
- Passion for mentorship and developing Production Engineering expertise

Apply on Crusoe's site: https://jobs.ashbyhq.com/Crusoe/d8019cfe-995a-40c3-bce0-97f368f3d454/application
