# Member of Technical Staff, Model Efficiency

Hiring organization: [Cohere](https://career.thegoodapps.co/organizations/cohere)

Canonical page: https://career.thegoodapps.co/jobs/b1a19df1-8c4a-4e4a-b763-ba344dd74541

Listed on Cohere's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Mid
- Location: New York, NY
- Remote: yes
- Salary: 250000 – 535000 CAD per year

## Summary

This role involves optimizing how large language models execute in production environments, focusing on reducing latency and increasing throughput through inference stack improvements. It suits experienced software engineers with strong systems programming skills who want to work on GPU optimization and performance engineering at scale.

_Our summary, not Cohere's wording._

## Skills named

C++, CUDA, Go, Python, Rust

## Required

- 5+ years writing high-performance production code
- Strong C++ or Python programming
- Experience with large language models and LLM inference ecosystem
- Ability to diagnose and resolve performance bottlenecks in model execution

## Nice to have

- GPU programming or CUDA experience
- Low-level systems optimization
- Language modeling with transformers including MoE and speculative decoding
- KV-cache optimization techniques
- Experience scaling performance-critical distributed systems

Apply on Cohere's site: https://jobs.ashbyhq.com/cohere/2a989030-6d14-4924-88c1-d878911e26fa/application
