# Staff Site Reliability Engineer, Ads

Hiring organization: [Reddit](https://career.thegoodapps.co/organizations/reddit)

Canonical page: https://career.thegoodapps.co/jobs/90d02d93-b138-483d-9e3a-aa1eeac8ab2e

Listed on Reddit's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Lead / Staff / Principal
- Location: San Francisco, CA
- Salary: 217000 – 303900 USD per year

## Summary

Lead reliability engineering for Reddit's ads systems and critical user-facing infrastructure at global scale. This role suits deeply technical engineers who excel in distributed systems, thrive on complex operational challenges, and can shape reliability practices across the organization.

_Our summary, not Reddit's wording._

## Skills named

Apache Cassandra, Apache Kafka, ClickHouse, Cloud Infrastructure, Go, Grafana, Kubernetes, OpenTelemetry, Prometheus, Python, Redis

## Required

- 8+ years in Site Reliability Engineering, Infrastructure Engineering, or equivalent roles with large-scale distributed systems
- Strong collaboration and communication skills with cross-team technical influence
- Experience supporting high-traffic user-facing production environments
- Deep understanding of distributed systems, networking, Linux systems, or cloud-native architectures
- Experience designing highly available systems with strong operational and reliability practices
- Strong programming skills in Go, Python, or similar languages
- Strong understanding of observability systems including metrics, logging, tracing, and alerting
- Experience improving reliability through SLOs, automation, incident management, and performance optimization
- Ability to troubleshoot complex issues across applications, infrastructure, networking, and services

## Nice to have

- Experience operating systems at internet-scale traffic volumes
- Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms
- Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies
- Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure
- Open source software contributions or participation in technical communities
- Experience leading large-scale incident response and operational transformation initiatives

Apply on Reddit's site: https://job-boards.greenhouse.io/reddit/jobs/7909463
