CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves optimizing how large language models execute in production environments, focusing on reducing latency and increasing throughput through inference stack improvements. It suits experienced software engineers with strong systems programming skills who want to work on GPU optimization and performance engineering at scale.
Our summary, not Cohere’s wording. The full posting is on their site.
What they ask for
Required
- 5+ years writing high-performance production code
- Strong C++ or Python programming
- Experience with large language models and LLM inference ecosystem
- Ability to diagnose and resolve performance bottlenecks in model execution
Nice to have
- GPU programming or CUDA experience
- Low-level systems optimization
- Language modeling with transformers including MoE and speculative decoding
- KV-cache optimization techniques
- Experience scaling performance-critical distributed systems