Lead Member of Technical Staff, Inference Infrastructure
CohereNew York, NY · Remote · full time · Lead
$235,000 – $325,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Lead the design and operation of high-performance ML infrastructure systems at an AI company, directing technical strategy for deploying large language models at scale across multiple cloud platforms. This role suits experienced infrastructure engineers who thrive in technical leadership, mentoring teams, and solving complex distributed systems challenges in production environments.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 8+ years production infrastructure engineering experience at scale
- Technical leadership track record
- Architecture and design of highly available distributed systems with Kubernetes and GPU workloads
- Deep Kubernetes expertise with production coding and support
- Experience across GCP, Azure, AWS, OCI, and multi-cloud/hybrid environments
- Design, deployment, and troubleshooting of large-scale Linux computing environments
- Ownership of compute, storage, and network resource management and cost optimization
- Mentoring and cross-functional leadership capabilities
- Understanding of GPU, TPU, and accelerator computational characteristics for latency and throughput optimization
- Distributed systems expertise with ability to establish team-wide patterns and practices
- Proficiency in Go, C++, or similar high-performance server languages