# Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

Hiring organization: [Crusoe](https://career.thegoodapps.co/organizations/crusoe)

Canonical page: https://career.thegoodapps.co/jobs/33e64dd4-bc49-4612-a90d-a6f0cadcd9b8

Listed on Crusoe's own careers site. Applications go to them directly.

- Seniority: Senior
- Location: San Francisco, CA
- Salary: 179000 – 218000 USD per year

## Summary

Crusoe seeks a senior technical authority to bridge GPU hardware architecture and data center operations, advising both engineering teams on next-generation facility design and operations teams on predictive maintenance and advanced troubleshooting for liquid-cooled AI infrastructure. This role suits someone with deep expertise in GPU platforms at hyperscale who can translate silicon roadmaps into operational strategy across thousands of nodes.

_Our summary, not Crusoe's wording._

## Skills named

Bash (Scripting), Go, Python

## Required

- 10+ years in hardware engineering, systems architecture, or data center infrastructure
- Expert knowledge of NVIDIA and AMD GPU architectures
- Ability to translate silicon datasheets into mechanical engineering requirements
- Proficiency in Python, Go, or Bash for telemetry and health-check tools
- Experience managing or architecting GPU clusters at scale at a hyperscaler or major silicon vendor
- Deep understanding of direct-to-chip cooling systems and thermal management
- B.S. or M.S. in Electrical Engineering, Computer Engineering, or related field

## Nice to have

- Experience with ML frameworks for predictive monitoring
- Vendor relationship management experience
- Root cause analysis on systemic hardware failures

Apply on Crusoe's site: https://jobs.ashbyhq.com/Crusoe/0a5354b7-a128-43a3-8f87-63a3109918b4/application
