Staff Research Engineer, Model Efficiency
CohereNew York, NY · Remote · full time · Lead / Staff / Principal
CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Cohere is seeking a Staff Research Engineer to develop and deploy techniques that improve the speed and efficiency of large language model inference in production. This role is ideal for someone with advanced ML expertise and a track record of publishing optimization research who wants to tackle the practical challenge of making AI systems run faster and more efficiently at scale.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- PhD in Machine Learning or related field
- Understanding of LLM architecture and inference optimization under resource constraints
- Significant experience with model efficiency enhancement techniques
- Strong software engineering skills
- Ability to work in fast-paced high-ambiguity startup environment
Nice to have
- Publications at top-tier conferences (ICLR, ACL, NeurIPS)
- Mentoring experience