CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role focuses on building and optimizing the infrastructure that trains Cohere's large language models at scale, working closely with researchers to improve training pipelines and performance. It suits engineers with strong software fundamentals and hands-on experience in distributed systems who want to bridge research and production in a compute-rich environment.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Extremely strong software engineering skills
- Proficiency in Python and ML frameworks (JAX, PyTorch, XLA/MLIR)
- Experience with distributed training infrastructures (Kubernetes, Slurm) and frameworks like Ray
- Experience with large-scale distributed training strategies
- Hands-on experience training large models at scale and contributing to training infrastructure tooling or setup
Nice to have
- Publication at top-tier ML and systems venues (NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)