Listed on CrowdStrike’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role leads reliability and scaling engineering for CrowdStrike's massive SIEM platform, handling petabytes of daily data and serving tens of thousands of customers. It suits experienced systems engineers with deep expertise in distributed systems, observability, and incident response who want to own end-to-end platform health across complex, interconnected pipelines.
Our summary, not CrowdStrike’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 10+ years in software engineering, SRE, or platform engineering
- Significant experience with large-scale distributed systems
- Proficiency in at least one systems language (Go, Java, Rust, C++)
- Proficiency in at least one scripting language (Python, Bash)
- Deep end-to-end observability experience
- Ability to diagnose and resolve complex multi-component incidents
- Experience with coordinated capacity planning and scaling
- Hands-on experience with Kafka or similar streaming platforms
- Experience with infrastructure-as-code and CI/CD
- Strong communication skills for incident leadership and post-mortem analysis
- Comfort working across time zones with distributed teams
Nice to have
- Experience at hyperscaler or large-scale SaaS provider
- Track record building automated remediation and self-healing systems
- Cost modeling experience for large compute and storage
- Cloud-native and serverless computing experience
- Experience operating platforms processing 1+ trillion events per day or 10+ PB daily
- Log management or cybersecurity product experience
- Disaster recovery planning for multi-region systems