# Senior Software Engineer, GPU Infrastructure (HPC)

Hiring organization: [Cohere](https://career.thegoodapps.co/organizations/cohere)

Canonical page: https://career.thegoodapps.co/jobs/51d9d309-29a1-4f12-94b1-c43dd27daf6e

Listed on Cohere's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Staff
- Location: Canada, United States
- Remote: yes
- Salary: 285000 – 340000 CAD per year

## Summary

This role builds and operates GPU/TPU infrastructure at scale for training frontier AI models, requiring deep expertise in Kubernetes, distributed systems, and close collaboration with AI researchers. It suits engineers with strong systems knowledge who thrive on solving complex infrastructure challenges in fast-paced environments.

_Our summary, not Cohere's wording._

## Skills named

Go, JAX, Kubernetes, Linux, Python, PyTorch, TensorFlow

## Required

- Deep expertise in ML/HPC infrastructure with GPU/TPU cluster experience
- Kubernetes deployment and management at scale for AI workloads
- Proficiency in Python and Go
- Experience with distributed training frameworks like JAX, PyTorch, or TensorFlow
- Linux internals and performance optimization knowledge
- Track record of collaborating with AI researchers on infrastructure challenges

## Nice to have

- RDMA networking expertise
- Open-source contributions
- Self-directed problem-solving in fast-paced environments
- Familiarity with infrastructure-as-code practices

Apply on Cohere's site: https://jobs.ashbyhq.com/cohere/ef9b939d-da66-464c-a878-ef45616c0473/application
