CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role optimizes the training performance of large language models by eliminating bottlenecks and writing low-level GPU kernels. It suits engineers with strong systems programming skills who want to work on the infrastructure that makes frontier AI models efficient.
Our summary, not Cohere’s wording. The full posting is on their site.
What they ask for
Required
- Extremely strong software engineering skills
- Proficiency in Python and ML frameworks (JAX, PyTorch, XLA/MLIR)
- Experience writing GPU kernels with CUDA or Triton
- Experience with large-scale distributed training
- Familiarity with autoregressive sequence models like Transformers
Nice to have
- Published papers at top-tier ML venues (NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)