Skip to main content
CareerApp

Site Reliability Engineer, Inference Infrastructure

Cohere

New York, NY · Remote · full time · Senior

$160,000 – $260,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role develops and operates the infrastructure that serves Cohere's language models at scale, focusing on Kubernetes automation, GPU clusters, and high-availability systems. It suits engineers with production infrastructure experience who want to work on the platform layer enabling enterprise AI deployment.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years production infrastructure engineering at scale
  • Kubernetes design and GPU cluster experience
  • Linux computing environment expertise
  • Distributed systems understanding
  • Compute and cost resource management

Nice to have

  • GCP, AWS, Azure, OCI or multi-cloud/hybrid serving experience
  • Golang, C++ or high-performance server language experience
  • Knowledge of GPU, TPU or custom accelerator characteristics
  • Familiarity with inference latency and throughput optimization

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.