Senior Site Reliability Engineer (Disaster Recovery at SunCore Digital | AZ, US | Rezi

Senior Site Reliability Engineer (Disaster Recovery at SunCore Digital

Senior Site Reliability Engineer (Disaster Recovery

SunCore Digital · AZ, US

1 weeks ago

Senior Site Reliability Engineer (Disaster Recovery

SunCore Digital · AZ, US

10 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this Senior Site Reliability Engineer (Disaster Recovery role.

Rezi rewrites your resume against SunCore Digital's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the Senior Site Reliability Engineer (Disaster Recovery posting at SunCore Digital — free, in seconds.

About the Role

SunCore Digital is seeking a hands-on Senior Site Reliability Engineer (Disaster Recovery) to assess, implement, test, and document service and data recovery capabilities. This is an engineering position focused on identifying recovery risks, building and hardening shared recovery capabilities, automating recovery processes, and proving critical systems and data can be restored.

Responsibilities

  • Inventory critical applications, services, databases, storage systems, queues, infrastructure, and external recovery dependencies.
  • Determine current backup status, including creation, retention, protection, and monitoring.
  • Identify systems with missing, incomplete, unverified, or person-dependent recovery processes.
  • Assess risks related to data loss, service loss, infrastructure failure, configuration loss, credential availability, and third-party dependencies.
  • Distinguish configured backups from proven recoverability.
  • Identify manual recovery procedures and undocumented knowledge.
  • Document confirmed capabilities, untested capabilities, and unknowns.
  • Design and implement backup improvements for critical data and configuration, then transfer ongoing operation to the designated long-term owner.
  • Translate approved recovery, retention, security, and access requirements into backup controls and monitoring.
  • Create, test, and harden restoration procedures with system owners, then transfer routine execution and system-specific maintenance to those teams.
  • Execute restoration tests using representative environments and data.
  • Validate backup completeness and integrity.
  • Measure recovery duration and potential data loss during exercises.
  • Establish monitoring and escalation for failed or incomplete backups.
  • Automate backup verification, restoration validation, recovery evidence collection, and other repeatable recovery processes, transferring routine operation to owning teams after testing and documentation.
  • Design and implement shared service-recovery and failover capabilities, transfer supporting shared components, and ensure system-owning teams retain responsibility for application-specific recovery behavior.
  • Identify infrastructure, application, database, network, access, and vendor dependencies required during recovery.
  • Establish safe recovery sequencing.
  • Develop procedures for partial outages, regional failures, infrastructure loss, data corruption, and service dependency failures.
  • Implement shared recovery mitigations and coordinate application-specific changes with system owners.
  • Test restored services for technical functionality.
  • Work with QA to validate critical business workflows and data remain correct after recovery.
  • Document conditions under which recovery, rollback, or failover may be unsafe.
  • Work with engineering and business leadership to document recovery requirements for critical systems.
  • Translate approved business requirements into technical recovery capabilities.
  • Measure actual recovery performance against defined expectations.
  • Identify where current architecture cannot meet required recovery expectations.
  • Provide technical options and evidence to support leadership decisions.
  • Maintain a recovery dependency map for critical systems.
  • Plan and execute recovery exercises.
  • Develop test scenarios for service, infrastructure, data, credential, and dependency failures.
  • Ensure exercises produce objective evidence.
  • Record results, defects, recovery times, data-loss observations, and unresolved risks.
  • Implement shared recovery corrective actions and coordinate system-specific corrective work with owning teams.
  • Repeat exercises after material changes.
  • Ensure procedures can be followed by qualified staff who did not author them.
  • Establish recovery runbook standards and create initial system-specific runbooks with owning teams.
  • Document required access, tools, credentials, dependencies, procedures, validation steps, and escalation paths.
  • Clearly label untested procedures.
  • Train system owners and relevant operational staff to execute and maintain system-specific procedures.
  • Reduce reliance on undocumented individual knowledge.
  • Ensure application teams understand their ongoing recovery responsibilities.
  • Maintain shared recovery evidence and exercise results.
  • Work with Security Operations to review recovery implementations affecting production access, sensitive data, credentials, networks, infrastructure, or security controls.
  • Work with SRE on service dependencies, failure behavior, observability, and operational response.
  • Define recovery requirements for infrastructure-as-code, environment recreation, artifacts, and delivery tooling.
  • Work with application engineers on application-specific recovery changes.
  • Work with QA on post-recovery functional and data validation.
  • Escalate recovery risks that cannot be mitigated within current architecture or resources.
  • Inventory existing backup and recovery capabilities.
  • Identify critical systems with no confirmed recovery path.
  • Verify ownership of each backup and recovery process.
  • Determine when critical backups were last restored successfully.
  • Establish repeatable restore testing.
  • Document initial recovery requirements and dependencies.
  • Implement the highest-priority recovery mitigations.
  • Create and test recovery runbooks.
  • Identify person-dependent and manual recovery processes.
  • Produce a factual recovery-readiness assessment before broader external use.

Requirements

  • Five or more years of experience in disaster recovery, infrastructure resilience, backup and recovery engineering, SRE, cloud infrastructure, systems engineering, or a related technical field.
  • Hands-on experience implementing backup, restore, recovery, and failover capabilities.
  • Experience performing recovery exercises rather than only writing recovery plans.
  • Experience recovering databases, cloud infrastructure, applications, configurations, and storage systems.
  • Experience with recovery automation and scripting.
  • Understanding of recovery-time and recovery-point concepts.
  • Experience documenting and validating recovery dependencies.
  • Familiarity with cloud security, identity, secrets, networking, and access requirements during recovery.
  • Experience troubleshooting complex system failures.
  • Ability to work with application developers on recovery-related code and data changes.
  • Strong runbook and technical-documentation skills.
  • Ability to identify unknowns and avoid treating untested procedures as proven.
  • Comfort working in a small organization where recovery practices are still being established.
  • Experience preparing a platform for its first external users.
  • Experience establishing a disaster-recovery capability from an early stage.
  • Experience with data-intensive, financial, digital-asset, telemetry, or operational systems.
  • Experience with infrastructure as code.
  • Experience with controlled disaster-recovery or resilience exercises.
  • Experience recovering event-driven or distributed systems.
  • Experience in a remote, asynchronous environment.

Skills

  • Disaster recovery
  • Infrastructure resilience
  • Backup and recovery engineering
  • SRE
  • Cloud infrastructure
  • Systems engineering
  • Backup implementation
  • Restore implementation
  • Recovery implementation
  • Failover implementation
  • Recovery exercises
  • Database recovery
  • Cloud infrastructure recovery
  • Application recovery
  • Configuration recovery
  • Storage system recovery
  • Recovery automation
  • Scripting
  • Recovery-time concepts
  • Recovery-point concepts
  • Recovery dependency documentation
  • Recovery dependency validation
  • Cloud security
  • Identity management during recovery
  • Secrets management during recovery
  • Networking during recovery
  • Access control during recovery
  • Troubleshooting complex system failures
  • Runbook creation
  • Technical documentation
  • Identifying unknowns
  • Establishing recovery practices
  • Platform preparation for external users
  • Disaster-recovery capability establishment
  • Data-intensive systems recovery
  • Financial systems recovery
  • Digital-asset systems recovery
  • Telemetry systems recovery
  • Operational systems recovery
  • Infrastructure as code
  • Controlled disaster-recovery exercises
  • Controlled resilience exercises
  • Event-driven systems recovery
  • Distributed systems recovery
  • Remote work
  • Asynchronous work

Location

  • Remote

Work Type

  • Fully remote
  • Async flexibility

Experience Level

  • Senior

Salary/Compensations

  • Competitive, based on experience, portfolio strength, and geographic location.
  • 10% milestone bonus awarded to every team member assisting with the MVP build-out.
  • Annual Performance Bonus: Represents a significant percentage of total compensation.

Benefits

  • Health, dental, and vision plans.
  • Annual Paid Offsite: Team retreats in Caribbean, Hawaii, ski destinations, and other exciting locations.

About the Company

  • High-impact role in a rapidly scaling digital company.
  • Fully remote team with async flexibility.
  • Direct collaboration with executive leadership and best-in-class marketing/design partners.
  • Opportunity to shape Security processes and mentor junior team members.
  • Competitive compensation with performance upside.