# Engineering Manager, GPU Infrastructure

Hiring organization: [Cohere](https://career.thegoodapps.co/organizations/cohere)

Canonical page: https://career.thegoodapps.co/jobs/bbde0f26-fd88-4b6c-b505-715ae3c1a68f

Listed on Cohere's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Senior
- Location: United States, Canada
- Salary: 330000 – 400000 CAD per year

## Summary

Lead a team of engineers building and operating the GPU clusters that train frontier AI models, setting technical direction for how these superclusters are deployed, scheduled, and scaled across multi-cloud environments. This role suits experienced infrastructure leaders with deep Kubernetes expertise who want to work at the intersection of hardware, distributed systems, and AI at a company training cutting-edge models.

_Our summary, not Cohere's wording._

## Skills named

Infrastructure as Code, Kubernetes

## Required

- Experience managing engineering or SRE teams with technical mentorship and hiring
- Running large Kubernetes compute fleets in production across multi-cloud environments
- Deep expertise in at least one layer of GPU training infrastructure (cluster operations, GPU networking, or hardware)
- Cost optimization and capacity planning for GPU infrastructure
- Track record partnering with researchers or ML engineers on reliability, cost, and delivery tradeoffs

## Nice to have

- Willingness to get hands-on and learn infrastructure layers beyond your specialty

Apply on Cohere's site: https://jobs.ashbyhq.com/cohere/28239d75-5dd9-41fb-ba43-cb08b491be2b/application
