$325,000 – $485,000
Listed on Menlo Ventures Portfolio’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role leads reliability engineering for Anthropic's AI serving infrastructure, designing monitoring systems and managing incidents across the critical path that delivers Claude to users. It suits engineers or SREs with distributed systems experience who can work across teams and think holistically about system resilience at scale.
Our summary, not Menlo Ventures Portfolio’s wording. The full posting is on their site.
What they ask for
Required
- Distributed systems, infrastructure, or reliability engineering background
- Strong communication and collaboration skills
- Ability to work across teams and build relationships
- Ownership mindset and care for user outcomes
Nice to have
- Prior SRE or Production Engineer role on large-scale systems
- Experience operating large-scale model serving or training infrastructure with >1000 GPUs
- Experience with ML hardware accelerators
- ML-specific networking optimization knowledge
- AI-specific observability tools expertise
- Chaos engineering and resilience testing experience
- Open-source infrastructure or ML tooling contributions