# Site Reliability Engineer, Inference Infrastructure

Hiring organization: [Cohere](https://career.thegoodapps.co/organizations/cohere)

Canonical page: https://career.thegoodapps.co/jobs/8ae8a3fc-f814-49db-a58b-74f6bcea6b50

Listed on Cohere's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Senior
- Location: New York, NY
- Remote: yes
- Salary: 160000 – 260000 USD per year

## Summary

This role develops and operates the infrastructure that serves Cohere's language models at scale, focusing on Kubernetes automation, GPU clusters, and high-availability systems. It suits engineers with production infrastructure experience who want to work on the platform layer enabling enterprise AI deployment.

_Our summary, not Cohere's wording._

## Skills named

Amazon Web Services (AWS), C++, Kubernetes, Linux, Microsoft Azure

## Required

- 5+ years production infrastructure engineering at scale
- Kubernetes design and GPU cluster experience
- Linux computing environment expertise
- Distributed systems understanding
- Compute and cost resource management

## Nice to have

- GCP, AWS, Azure, OCI or multi-cloud/hybrid serving experience
- Golang, C++ or high-performance server language experience
- Knowledge of GPU, TPU or custom accelerator characteristics
- Familiarity with inference latency and throughput optimization

Apply on Cohere's site: https://jobs.ashbyhq.com/cohere/8b6696e1-f1c4-4010-bde9-3cec1340a2a6/application
