Skip to main content
CareerApp

Audio Inference Engineer, Model Efficiency

Cohere

New York, NY · Remote · full time

CA$250,000 – CA$535,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Cohere seeks an engineer to optimize audio model inference performance across latency, throughput, and quality metrics. This role suits someone with deep systems expertise in machine learning inference who can identify bottlenecks and deliver solutions for real-time audio processing at scale.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • High-performance audio or machine learning inference systems development
  • C++ and Python proficiency
  • Deep learning models for audio, speech, or language applications
  • Results-oriented mindset

Nice to have

  • GPU programming and low-level system optimization
  • Model parallelization across multiple GPUs
  • Duplex real-time streaming architectures
  • Machine learning framework internals for audio
  • Custom distributed inference systems
  • Transformers and sequence modeling for audio/speech
  • End-to-end audio pipeline optimization

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.