Skip to main content
CareerApp

Staff Production Engineer, Core PE

Crusoe

Sunnyvale, CA · full time · Staff

$209,000 – $253,000

Listed on Crusoe’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Crusoe seeks a senior production engineer to lead reliability and operational excellence across their GPU cloud infrastructure that powers AI workloads. This role combines technical depth in distributed systems and infrastructure with leadership responsibilities, mentoring junior engineers while architecting observability and automation solutions for large-scale, latency-sensitive environments.

Our summary, not Crusoe’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Bachelor's degree in Computer Science, Engineering, or related technical field or equivalent practical experience
  • 8+ years of Production Engineering, SRE, or large-scale infrastructure operations experience
  • Experience supporting GPU workloads, HPC environments, or latency/throughput-sensitive distributed systems
  • Previous experience in infrastructure roles building or managing compute, storage or networking platforms
  • Deep knowledge of Linux/Unix systems including kernel and user space debugging
  • Understanding of modern cloud infrastructure fundamentals including Kubernetes, distributed systems, virtualization, and cloud platforms
  • Track record with incident management practices and reliability frameworks
  • Hands-on experience with Prometheus and Grafana
  • Experience with infrastructure-as-code and configuration management tools
  • Proficiency in scripting or programming with Go, Python, C, or C++
  • Exceptional communication skills and cross-team collaboration ability
  • Ability to remain effective while troubleshooting complex production issues
  • Commitment to reliability engineering and operational excellence

Nice to have

  • Experience leading Kubernetes or container orchestration platforms at scale
  • Exposure to change management processes, operational readiness reviews, or structured root cause analysis
  • Experience designing self-healing systems, automated remediation, or event-driven operational tooling
  • Interest in scaling AI or HPC infrastructure and GPU-heavy reliability challenges
  • Passion for mentorship and developing Production Engineering expertise

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.