Skip to main content
CareerApp

Engineering Manager, GPU Infrastructure

Cohere

United States, Canada · full time · Senior

CA$330,000 – CA$400,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Lead a team of engineers building and operating the GPU clusters that train frontier AI models, setting technical direction for how these superclusters are deployed, scheduled, and scaled across multi-cloud environments. This role suits experienced infrastructure leaders with deep Kubernetes expertise who want to work at the intersection of hardware, distributed systems, and AI at a company training cutting-edge models.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Experience managing engineering or SRE teams with technical mentorship and hiring
  • Running large Kubernetes compute fleets in production across multi-cloud environments
  • Deep expertise in at least one layer of GPU training infrastructure (cluster operations, GPU networking, or hardware)
  • Cost optimization and capacity planning for GPU infrastructure
  • Track record partnering with researchers or ML engineers on reliability, cost, and delivery tradeoffs

Nice to have

  • Willingness to get hands-on and learn infrastructure layers beyond your specialty

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.