About the Role
Join the NG-SIEM EPICS team as a Senior Engineer II to own the reliability and scalability of the security industry's largest SIEM platform. This role focuses on solving reliability and scalability as software engineering problems, building observability, automation, and scaling systems for the entire platform, which processes over 100 PB of data daily with up to 10 years of retention and millions of queries per hour.
Responsibilities
- Design, build, and maintain end-to-end observability monitoring and synthetic test suites for the NG-SIEM pipeline.
- Engineer orchestrated scaling solutions for the NG-SIEM pipeline as a unified system.
- Serve as a subject matter expert during platform-wide incidents, diagnosing and resolving multi-component failures.
- Participate in follow-the-sun on-call rotations, coordinating critical platform-wide events.
- Build and refine models for end-to-end capacity forecasting.
- Develop tooling to continuously track and surface cost drivers across the platform.
- Transform manual standard operating procedures into automated remediation workflows.
- Partner with various teams and stakeholders to triage SLO breaches and drive problem management.
- Identify and drive systemic platform improvements for long-term resilience and efficiency.
Requirements
- Passion for reliability engineering and curiosity about large-scale system behavior under pressure.
- 10+ years of experience in software engineering, site reliability engineering, or platform engineering, with large-scale distributed systems.
- Ability to make pragmatic tradeoffs between short-term delivery and long-term platform goals.
- Strong proficiency in at least one systems programming language (Go, Java, Rust, or C++).
- Strong proficiency in at least one scripting language (Python, Bash).
- Deep experience with end-to-end observability, including monitoring pipelines, SLIs/SLOs, and dashboards.
- Demonstrated ability to diagnose and resolve complex incidents across multiple distributed components 24/7.
- Experience with coordinated capacity planning and scaling for significant infrastructure footprints.
- Hands-on experience with streaming platforms (Kafka or similar), understanding back pressure, partition management, and consumer group dynamics.
- Familiarity with infrastructure-as-code, CI/CD pipelines, and automated deployment practices.
- Collaborative team player with a proactive attitude.
- Strong written and verbal communication skills for incident leadership and post-incident analyses.
- Comfort working across time zones with globally distributed teams.
- Experience in a similar reliability or platform engineering role at a hyperscaler (AWS, Azure, GCP) or large-scale SaaS provider (Bonus).
- Track record of building automated remediation and self-healing infrastructure (Bonus).
- Experience with cost modeling and unit economics for large compute and storage footprints (Bonus).
- Familiarity with cloud-native architectures and serverless computing paradigms (Bonus).
- Hands-on experience operating platforms processing over 1 trillion events per day or more than 10 PB of data per day (Bonus).
- Exposure to or experience with Log Management, cybersecurity products, or security operations workflows (Bonus).
- Experience with disaster recovery planning and execution for multi-region systems (Bonus).
- Ability to periodically undergo and pass additional background and fingerprint check(s) consistent with government customer requirements.
Skills
- Go
- Java
- Rust
- C++
- Python
- Bash
- Reliability Engineering
- Site Reliability Engineering
- Platform Engineering
- Distributed Systems
- Observability
- Monitoring Pipelines
- SLIs/SLOs
- Dashboards
- Incident Response
- Capacity Planning
- Scaling
- Streaming Platforms
- Kafka
- Infrastructure-as-code
- CI/CD Pipelines
- Automated Deployment
- Automated Remediation
- Self-healing Infrastructure
- Cost Modeling
- Unit Economics
- Cloud-native Architectures
- Serverless Computing
- Log Management
- Cybersecurity
- Security Operations
- Disaster Recovery
Location
- Austin, TX
- Hybrid
Work Type
- Hybrid
- Full-time
Experience Level
- Senior
- 10+ years of experience
Salary/Compensations
- $160,000 - $250,000 per year (base salary)
- Eligibility for bonuses
- Eligibility for equity grants
Benefits
- Market leader in compensation and equity awards
- Comprehensive physical and mental wellness programs
- Competitive vacation and holidays
- Paid parental and adoption leaves
- Professional development opportunities
- Employee Networks
- Geographic neighborhood groups
- Volunteer opportunities
- Vibrant office culture with world-class amenities
- Great Place to Work Certified™
- Health insurance
- 401k
- Paid time off
About the Company
- Global leader in cybersecurity, protecting people, processes, and technologies for modern organizations.
- Mission since 2011: stop breaches, redefining modern security with the world’s most advanced AI-native platform.
- Works on large-scale distributed systems, processing almost 3 trillion events per day.
- Serves customers across all industries, ensuring business continuity and safety.
- Mission-driven company cultivating a culture of flexibility and autonomy for employees.
- Founded in 2011 to address sophisticated attacks unsolvable by malware-based defenses.
- Combines advanced endpoint protection with expert intelligence to identify adversaries.
Equal Opportunity
- CrowdStrike is an equal opportunity employer, committed to fostering a culture of belonging.
- Supports veterans and individuals with disabilities through an affirmative action program.
- Provides equal employment opportunity for all employees and applicants, without discrimination based on race, color, creed, ethnicity, religion, sex, sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, or any other characteristic protected by law.
- Bases all employment decisions on valid job requirements.
- Offers assistance for accessing information, submitting applications, or requesting accommodations via recruiting@crowdstrike.com.
- Participates in the E-Verify program.
