$195,000 – $235,000
Listed on Crusoe’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Crusoe is seeking a seasoned network operations engineer to ensure reliability across their global AI infrastructure, including GPU cluster interconnects and data center networks. This role combines hands-on incident response, automation development, and operational leadership for someone who thrives under pressure and wants to directly impact hyperscale AI availability.
Our summary, not Crusoe’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 8+ years of production network engineering in large-scale environments
- Strong Python and scripting proficiency for diagnostic and auto-remediation tooling
- Experience with observability tools including streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes
- Hands-on experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads including PFC, ECN, and DCQCN tuning
- Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments
- Proficiency with Arista EOS and Juniper Junos in leaf-spine CLOS architectures
- Experience operating large device fleets across multi-region environments with on-call responsibility
- Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience
Nice to have
- Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility
- Experience defining or contributing to SLIs and SLOs in partnership with SRE or product teams
- Exposure to operating 10K+ device fleets across hyperscale or cloud environments
- Background contributing to post-incident learning programs or operational excellence initiatives