# Staff Network Production Engineer, Operations

Hiring organization: [Crusoe](https://career.thegoodapps.co/organizations/crusoe)

Canonical page: https://career.thegoodapps.co/jobs/a19213d0-64f5-4c5b-aeeb-d1bb5749771c

Listed on Crusoe's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Staff
- Location: San Francisco, CA
- Salary: 195000 – 235000 USD per year

## Summary

Crusoe is seeking a seasoned network operations engineer to ensure reliability across their global AI infrastructure, including GPU cluster interconnects and data center networks. This role combines hands-on incident response, automation development, and operational leadership for someone who thrives under pressure and wants to directly impact hyperscale AI availability.

_Our summary, not Crusoe's wording._

## Skills named

Border Gateway Protocol (BGP), Grafana, MPLS, OSPF, Prometheus, Python, TCP/IP

## Required

- 8+ years of production network engineering in large-scale environments
- Strong Python and scripting proficiency for diagnostic and auto-remediation tooling
- Experience with observability tools including streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, and ThousandEyes
- Hands-on experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads including PFC, ECN, and DCQCN tuning
- Expert knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments
- Proficiency with Arista EOS and Juniper Junos in leaf-spine CLOS architectures
- Experience operating large device fleets across multi-region environments with on-call responsibility
- Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience

## Nice to have

- Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility
- Experience defining or contributing to SLIs and SLOs in partnership with SRE or product teams
- Exposure to operating 10K+ device fleets across hyperscale or cloud environments
- Background contributing to post-incident learning programs or operational excellence initiatives

Apply on Crusoe's site: https://jobs.ashbyhq.com/Crusoe/ef99771f-c1a6-44eb-94f3-fae7652d7a47/application
