Skip to main content
CareerApp

Staff Software Engineer, AI Reliability

Menlo Ventures Portfolio

New York, NY · full time · Staff

$325,000 – $485,000

Listed on Menlo Ventures Portfolio’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role leads reliability engineering for Anthropic's AI serving infrastructure, designing monitoring systems and managing incidents across the critical path that delivers Claude to users. It suits engineers or SREs with distributed systems experience who can work across teams and think holistically about system resilience at scale.

Our summary, not Menlo Ventures Portfolio’s wording. The full posting is on their site.

What they ask for

Required

  • Distributed systems, infrastructure, or reliability engineering background
  • Strong communication and collaboration skills
  • Ability to work across teams and build relationships
  • Ownership mindset and care for user outcomes

Nice to have

  • Prior SRE or Production Engineer role on large-scale systems
  • Experience operating large-scale model serving or training infrastructure with >1000 GPUs
  • Experience with ML hardware accelerators
  • ML-specific networking optimization knowledge
  • AI-specific observability tools expertise
  • Chaos engineering and resilience testing experience
  • Open-source infrastructure or ML tooling contributions

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.