About the Role
This role involves leading the architecture of highly scalable, interdependent distributed systems, identifying and removing performance bottlenecks, and designing elastic, high-impact systems. The position also focuses on advancing innovation in data plane platforms, engineering fault-tolerant designs, optimizing resilience, and setting availability standards.
Responsibilities
- Mentors teams and leads the architecture of highly scalable, interdependent distributed systems.
- Identifies and removes performance/scalability bottlenecks for hyper‑scale workloads.
- Defines scalability requirements with stakeholders.
- Designs elastic, high‑impact systems while advancing innovation in data plane platforms.
- Engineers and oversees fault‑tolerant, in‑service‑upgradable designs.
- Optimizes resilience mechanisms (load‑shedding, throttling, rate‑limiting).
- Sets SLO‑aligned durability and availability standards across dependent services.
- Establishes KPIs and advanced telemetry.
- Applies formal verification for complex features.
- Develops robust replication/synchronization strategies.
- Advises and leads resolution of complex production issues.
- Sets operational readiness and SOP standards.
- Directs incident response and RCAs.
- Architects advanced security controls.
- Drives remediation and compliance.
- Delivers enterprise‑level automation (IaC) and change strategies enabling safe, automated patching, updates, and rollbacks.
Requirements
- Experience with hyper‑scale workloads.
- Experience with fault‑tolerant, in‑service‑upgradable designs.
- Experience with resilience mechanisms (load‑shedding, throttling, rate‑limiting).
- Experience with SLO‑aligned durability and availability standards.
- Experience with KPIs and advanced telemetry.
- Experience with formal verification for complex features.
- Experience with replication/synchronization strategies.
- Experience with production issue resolution.
- Experience with operational readiness and SOP standards.
- Experience with incident response and RCAs.
- Experience with advanced security controls.
- Experience with remediation and compliance.
- Experience with enterprise‑level automation (IaC) and change strategies.
Skills
- Architecture
- Scalability
- Performance Optimization
- Distributed Systems
- Data Plane Platforms
- Fault Tolerance
- Resilience Engineering
- Service Level Objectives (SLOs)
- Key Performance Indicators (KPIs)
- Telemetry
- Formal Verification
- Replication
- Synchronization
- Incident Response
- Root Cause Analysis (RCA)
- Security Controls
- Infrastructure as Code (IaC)
- Automation
- Change Management
Benefits
- Flexible medical insurance
- Life insurance
- Retirement options
- Volunteer programs
About the Company
- Oracle brings together data, infrastructure, applications, and expertise to power industry innovations and life-saving care.
- Oracle embeds AI across its products and services to help customers turn promise into a better future.
- Oracle is a company leading the way in AI and cloud solutions that impact billions of lives.
Equal Opportunity
- Oracle is committed to growing a workforce that promotes opportunities for all.
- We are committed to including people with disabilities at all stages of the employment process.
- Oracle is an Equal Employment Opportunity Employer.
- All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law.
- Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
