CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Cohere seeks a senior engineer to design and maintain the training framework for frontier-scale language models, focusing on distributed systems, HPC infrastructure, and the tooling that connects research to thousands of GPUs. This role suits someone with deep experience in large-scale distributed training who enjoys working across the full ML systems stack to solve performance and reliability challenges.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong engineering experience in large-scale distributed training or HPC systems
- Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops
- Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)
- Ability to debug performance issues across CUDA/NCCL, networking, IO, and data pipelines
- Experience with containerized environments (Docker, Singularity/Apptainer)
- Track record of building tools that increase developer velocity for ML teams
- Sound judgment balancing performance vs complexity and research velocity vs maintainability
- Strong collaboration skills across infrastructure, research, and deployment teams
Nice to have
- Experience training LLMs or other large transformer architectures
- Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers)
- Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches)
- Experience with data pipeline optimization, sharded datasets, or caching strategies
- Background in performance engineering, profiling, or low-level systems
- Publication record at top-tier venues (NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)