Skip to main content
CareerApp

Senior Software Engineer, GPU Infrastructure (HPC)

Cohere

Canada, United States · Remote · full time · Staff

CA$285,000 – CA$340,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role builds and operates GPU/TPU infrastructure at scale for training frontier AI models, requiring deep expertise in Kubernetes, distributed systems, and close collaboration with AI researchers. It suits engineers with strong systems knowledge who thrive on solving complex infrastructure challenges in fast-paced environments.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Deep expertise in ML/HPC infrastructure with GPU/TPU cluster experience
  • Kubernetes deployment and management at scale for AI workloads
  • Proficiency in Python and Go
  • Experience with distributed training frameworks like JAX, PyTorch, or TensorFlow
  • Linux internals and performance optimization knowledge
  • Track record of collaborating with AI researchers on infrastructure challenges

Nice to have

  • RDMA networking expertise
  • Open-source contributions
  • Self-directed problem-solving in fast-paced environments
  • Familiarity with infrastructure-as-code practices

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.