Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.
13 open listings
Roles posted by organizations here, alongside roles we found on employers’ own careers sites. Crawled roles say so on the card and send you to the employer to apply. Only show organizations on Career App
LangChain
Full-time · San Francisco, CA · $150,000 – $190,000
You'll own the health and reliability of LangChain's production AI agent that supports the go-to-market team, monitoring its performance, costs, and business impact while building the feedback loops and monitoring systems that keep it improving. This role suits engineers with strong production Python and SQL skills, real experience running LLM applications, and SRE instincts who want to both operate a complex system and evangelize the practices they build.
Listed on LangChain’s careers site · Apply there ↗
Shield AI
Full-time · San Diego, CA · $190,000 – $280,000
Shield AI seeks an experienced SRE leader to establish and mature reliability practices across its cloud infrastructure and platform services. The role combines hands-on technical work—investigating failures, improving tooling, and driving incident learnings—with mentorship and strategic leadership to build a reliability-focused culture.
Listed on Shield AI’s careers site · Apply there ↗
Full-time · San Francisco, CA · $190,800 – $267,100
This role manages the reliability and performance of Reddit's advertising platform systems, partnering with engineering teams to build infrastructure that serves high-traffic ad-delivery, auction, and billing systems. It suits experienced infrastructure engineers who want to own critical operational systems at scale and drive improvements across a revenue-critical product.
Listed on Reddit’s careers site · Apply there ↗
Anyscale
San Francisco, CA · $200,000 – $240,000 · Remote
This infrastructure engineering role focuses on designing and optimizing the distributed systems that power Anyscale's Ray platform, handling everything from Kubernetes orchestration to GPU integration. It suits experienced engineers with deep cloud infrastructure expertise who want to work on both open-source projects and production systems serving AI applications at scale.
Listed on Anyscale’s careers site · Apply there ↗
Circle
San Francisco, CA · $152,500 – $205,000 · Remote
Circle is hiring a Senior Site Reliability Engineer to design and operate secure, scalable platform infrastructure for digital assets and blockchain workloads across hybrid and public cloud. This role suits experienced infrastructure engineers who want to own production reliability, build Kubernetes platforms, and mentor teams while working on high-stakes financial systems.
Listed on Circle’s careers site · Apply there ↗
Crusoe
Full-time · Sunnyvale, CA · $170,000 – $205,000
Crusoe seeks a Storage SRE to build and maintain distributed cloud storage infrastructure supporting AI workloads, with a focus on automation, reliability, and performance optimization across block, file, and object storage systems. The role suits experienced storage engineers who want to work on mission-critical infrastructure at a vertically integrated AI company.
Listed on Crusoe’s careers site · Apply there ↗
Intuit
2 locations · $202,500 – $274,000
This is a senior site reliability engineering role focused on maintaining highly available fintech infrastructure serving millions of small business users. The position combines hands-on systems design, automation, and incident leadership within a distributed cloud environment at scale.
Listed on Intuit’s careers site · Apply there ↗
Tatari
Full-time · 3 locations · $190,000 – $240,000
This is a systems and infrastructure role focused on keeping Tatari's data platform reliable and stable across all environments, rather than building data pipelines. It suits experienced SRE or platform engineers who have learned to respect production through hard experience and bring operational discipline to infrastructure scaling and deployment.
Listed on Tatari’s careers site · Apply there ↗
Bend Studio
Full-time · Aliso Viejo, CA · $145,700 – $218,500
A Site Reliability Engineer focused on cloud gaming infrastructure, managing API gateways, service meshes, and distributed systems at scale for PlayStation's gaming platform. This role suits engineers with strong Linux and production systems experience who want to own reliability and operational excellence across a high-traffic gaming service.
Listed on Bend Studio’s careers site · Apply there ↗
Shield AI
Full-time · San Mateo, CA · $220,000 – $339,974
Lead the establishment and maturation of Site Reliability Engineering practices across Shield AI's cloud infrastructure and platform services. This hands-on technical role combines incident investigation, systems improvement, and mentorship to build a reliability-first engineering culture.
Listed on Shield AI’s careers site · Apply there ↗
Circle
San Francisco, CA · $195,000 – $257,500 · Remote
Circle is hiring a Staff Site Reliability Engineer to design, build, and operate the distributed infrastructure powering their blockchain platform at scale. This role suits experienced engineers who excel at managing complex systems, automating operations, and mentoring teams while working across multiple cloud environments and blockchain networks.
Listed on Circle’s careers site · Apply there ↗
Full-time · San Francisco, CA · $217,000 – $303,900
Lead reliability engineering for Reddit's ads systems and critical user-facing infrastructure at global scale. This role suits deeply technical engineers who excel in distributed systems, thrive on complex operational challenges, and can shape reliability practices across the organization.
Listed on Reddit’s careers site · Apply there ↗
Crusoe
San Francisco, CA · $156,000 – $190,000
Crusoe seeks a Staff-level cloud infrastructure expert to lead technical escalations, design reliability systems, and mentor engineers across their AI compute platform. This role suits experienced SRE or DevOps engineers who want to shape infrastructure architecture while directly supporting high-stakes customer incidents in a fast-growing AI infrastructure company.
Listed on Crusoe’s careers site · Apply there ↗