$217,000 – $303,900
Listed on Reddit’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Lead reliability engineering for Reddit's ads systems and critical user-facing infrastructure at global scale. This role suits deeply technical engineers who excel in distributed systems, thrive on complex operational challenges, and can shape reliability practices across the organization.
Our summary, not Reddit’s wording. The full posting is on their site.
Skills this role names
- Apache Cassandra
- Apache Kafka
- ClickHouse
- Cloud Infrastructure
- Go
- Grafana
- Kubernetes
- OpenTelemetry
- Prometheus
- Python
- Redis
Log in to see which of these are already on your profile.
What they ask for
Required
- 8+ years in Site Reliability Engineering, Infrastructure Engineering, or equivalent roles with large-scale distributed systems
- Strong collaboration and communication skills with cross-team technical influence
- Experience supporting high-traffic user-facing production environments
- Deep understanding of distributed systems, networking, Linux systems, or cloud-native architectures
- Experience designing highly available systems with strong operational and reliability practices
- Strong programming skills in Go, Python, or similar languages
- Strong understanding of observability systems including metrics, logging, tracing, and alerting
- Experience improving reliability through SLOs, automation, incident management, and performance optimization
- Ability to troubleshoot complex issues across applications, infrastructure, networking, and services
Nice to have
- Experience operating systems at internet-scale traffic volumes
- Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms
- Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies
- Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure
- Open source software contributions or participation in technical communities
- Experience leading large-scale incident response and operational transformation initiatives