Skip to main content
CareerApp

Staff Research Engineer, Model Efficiency

Cohere

New York, NY · Remote · full time · Lead / Staff / Principal

CA$250,000 – CA$535,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Cohere is seeking a Staff Research Engineer to develop and deploy techniques that improve the speed and efficiency of large language model inference in production. This role is ideal for someone with advanced ML expertise and a track record of publishing optimization research who wants to tackle the practical challenge of making AI systems run faster and more efficiently at scale.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • PhD in Machine Learning or related field
  • Understanding of LLM architecture and inference optimization under resource constraints
  • Significant experience with model efficiency enhancement techniques
  • Strong software engineering skills
  • Ability to work in fast-paced high-ambiguity startup environment

Nice to have

  • Publications at top-tier conferences (ICLR, ACL, NeurIPS)
  • Mentoring experience

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.