# Site Reliability Engineer

Hiring organization: [Cognition AI](https://career.thegoodapps.co/organizations/cognition-ai)

Canonical page: https://career.thegoodapps.co/jobs/52d75ef4-2dc5-4148-8862-81edff26aa42

Listed on Cognition AI's own careers site. Applications go to them directly.

- Employment type: full time
- Location: New York, NY
- Salary: 260000 – 300000 USD per year

## Summary

This role owns production reliability and platform engineering for Devin and Windsurf, AI developer tools used by hundreds of thousands daily. You'll define SLOs, lead incident response, build CI/CD pipelines, and partner with product teams to engineer reliability from the start—combining on-call ownership with infrastructure-as-code and observability work.

_Our summary, not Cognition AI's wording._

## Skills named

Amazon Web Services (AWS), CI/CD, Infrastructure as Code, Kubernetes, Microsoft Azure, Terraform

## Required

- Production systems experience at scale with SLOs and error budgets
- Strong software engineering fundamentals
- Cloud infrastructure proficiency
- Container orchestration experience
- Incident response and on-call ownership
- Observability and monitoring expertise
- Systematic toil reduction through automation

## Nice to have

- Experience with developer-facing products or platforms

Apply on Cognition AI's site: https://jobs.ashbyhq.com/cognition/d50d94b0-60c8-4dae-9c36-234f072ee4e3/application
