Skip to main content
CareerApp

Member of Technical Staff, Model Efficiency

Cohere

New York, NY · Remote · full time · Mid

CA$250,000 – CA$535,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role involves optimizing how large language models execute in production environments, focusing on reducing latency and increasing throughput through inference stack improvements. It suits experienced software engineers with strong systems programming skills who want to work on GPU optimization and performance engineering at scale.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years writing high-performance production code
  • Strong C++ or Python programming
  • Experience with large language models and LLM inference ecosystem
  • Ability to diagnose and resolve performance bottlenecks in model execution

Nice to have

  • GPU programming or CUDA experience
  • Low-level systems optimization
  • Language modeling with transformers including MoE and speculative decoding
  • KV-cache optimization techniques
  • Experience scaling performance-critical distributed systems

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.