Rodeo
Get started

SunCore Digital

Senior Site Reliability Engineer (Disaster Recovery

London
$145k – $200k/yr
Posted about 2 months ago
Sign up to applySee more jobs like this
Get notified of more jobs like this · No spam, ever

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Senior Site Reliability Engineer (Disaster Recovery)

About the Role

SunCore Digital is seeking a hands-on Senior Site Reliability Engineer (Disaster Recovery) to assess, implement, test, and document service and data recovery capabilities.

This is an engineering position, not solely a policy, compliance, or coordination role. The successful candidate will identify recovery risks, build and harden shared recovery capabilities, automate recovery-specific processes where practical, and prove that critical systems and data can be restored. Once system-specific procedures are tested and documented, the owning teams will be trained to execute and maintain them.

The engineer may draw on application developers when recovery improvements require changes to application code, databases, data flows, or system-specific behavior. Recovery procedures must ultimately be understood and maintainable by the teams that own the affected systems.

Responsibilities

Recovery assessment

  • Inventory critical applications, services, databases, storage systems, queues, infrastructure, and external recovery dependencies.
  • Determine what is currently backed up and how those backups are created, retained, protected, and monitored.
  • Identify systems with missing, incomplete, unverified, or person-dependent recovery processes.
  • Assess risks related to data loss, service loss, infrastructure failure, configuration loss, credential availability, and third-party dependencies.
  • Distinguish configured backups from proven recoverability.
  • Identify manual recovery procedures and undocumented knowledge.
  • Document confirmed capabilities, untested capabilities, and unknowns.

Backup and restoration

  • Design and implement backup improvements for critical data and configuration, then transfer ongoing operation to the designated long-term owner.
  • Translate approved recovery, retention, security, and access requirements into backup controls and monitoring.
  • Create, test, and harden restoration procedures with system owners, then transfer routine execution and system-specific maintenance to those teams.
  • Execute restoration tests using representative environments and data.
  • Validate backup completeness and integrity.
  • Measure recovery duration and potential data loss during exercises.
  • Establish monitoring and escalation so failed or incomplete backups are detected by the accountable long-term owner.
  • Automate backup verification, restoration validation, recovery evidence collection, and other repeatable recovery processes where practical; transfer routine operation of system-specific automation to the owning teams after it is tested and documented.

A backup will not be treated as reliable solely because a scheduled job reports success. Restoration must be tested.

Service recovery and failover

  • Design and implement shared service-recovery and failover capabilities, transfer supporting shared components to the designated long-term owner, and ensure system-owning teams retain responsibility for application-specific recovery behavior and procedures after handoff.
  • Identify the infrastructure, application, database, network, access, and vendor dependencies required during recovery.
  • Establish safe recovery sequencing.
  • Develop procedures for partial outages, regional failures, infrastructure loss, data corruption, and service dependency failures.
  • Implement shared recovery mitigations directly and coordinate application-specific changes with system owners, who retain responsibility for those changes.
  • Test restored services for technical functionality.
  • Work with QA to validate that critical business workflows and data remain correct after recovery.
  • Document conditions under which recovery, rollback, or failover may be unsafe.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Recovery objectives and planning

  • Work with engineering and business leadership to document recovery requirements for critical systems.
  • Translate approved business requirements into technical recovery capabilities.
  • Measure actual recovery performance against defined expectations.
  • Identify where current architecture cannot meet required recovery expectations.
  • Provide technical options and evidence to support leadership decisions.
  • Maintain a recovery dependency map for critical systems.

Final recovery priorities, risk acceptance, and business requirements remain with leadership.

Disaster-recovery exercises

  • Plan and execute recovery exercises.
  • Develop test scenarios for service, infrastructure, data, credential, and dependency failures.
  • Ensure exercises produce objective evidence.
  • Record results, defects, recovery times, data-loss observations, and unresolved risks.
  • Implement shared recovery corrective actions and coordinate system-specific corrective work with the owning teams.
  • Repeat exercises after material changes.
  • Ensure procedures can be followed by qualified staff who did not author them.

Documentation and knowledge transfer

  • Establish recovery runbook standards and create initial system-specific runbooks with the owning teams.
  • Document required access, tools, credentials, dependencies, procedures, validation steps, and escalation paths.
  • Clearly label untested procedures.
  • Train system owners and relevant operational staff to execute and maintain the system-specific procedures they own.
  • Reduce reliance on undocumented individual knowledge.
  • Ensure application teams understand their ongoing recovery responsibilities.
  • Maintain shared recovery evidence and exercise results; system owners maintain the accuracy of their service-specific runbooks.

Security and collaboration

  • Work with Security Operations to review recovery implementations that affect production access, sensitive data, credentials, networks, infrastructure, or security controls.
  • Work with SRE on service dependencies, failure behavior, observability, and operational response.
  • Define recovery requirements for infrastructure-as-code, environment recreation, artifacts, and delivery tooling; work with DevOps and platform staff to implement those requirements within their shared capabilities.
  • Work with application engineers on application-specific recovery changes.
  • Work with QA on post-recovery functional and data validation.
  • Escalate recovery risks that cannot be mitigated within current architecture or resources.

Initial Priorities

  • Inventory existing backup and recovery capabilities.
  • Identify critical systems with no confirmed recovery path.
  • Verify who owns each backup and recovery process.
  • Determine when critical backups were last restored successfully.
  • Establish repeatable restore testing.
  • Document initial recovery requirements and dependencies.
  • Implement the highest-priority recovery mitigations.
  • Create and test recovery runbooks.
  • Identify person-dependent and manual recovery processes.
  • Produce a factual recovery-readiness assessment before broader external use.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Required Qualifications

  • Five or more years of experience in disaster recovery, infrastructure resilience, backup and recovery engineering, SRE, cloud infrastructure, systems engineering, or a related technical field.
  • Hands-on experience implementing backup, restore, recovery, and failover capabilities.
  • Experience performing recovery exercises rather than only writing recovery plans.
  • Experience recovering databases, cloud infrastructure, applications, configurations, and storage systems.
  • Experience with recovery automation and scripting.
  • Understanding of recovery-time and recovery-point concepts.
  • Experience documenting and validating recovery dependencies.
  • Familiarity with cloud security, identity, secrets, networking, and access requirements during recovery.
  • Experience troubleshooting complex system failures.
  • Ability to work with application developers on recovery-related code and data changes.
  • Strong runbook and technical-documentation skills.
  • Ability to identify unknowns and avoid treating untested procedures as proven.
  • Comfort working in a small organization where recovery practices are still being established.

Bonus Points

  • Experience preparing a platform for its first external users.
  • Experience establishing a disaster-recovery capability from an early stage.
  • Experience with data-intensive, financial, digital-asset, telemetry, or operational systems.
  • Experience with infrastructure as code.
  • Experience with controlled disaster-recovery or resilience exercises.
  • Experience recovering event-driven or distributed systems.
  • Experience in a remote, asynchronous environment.

What Success Looks Like

  • Critical systems and data have documented recovery requirements.
  • Backup ownership, schedules, retention, and locations are known.
  • Critical backups have been restored successfully.
  • Recovery procedures are documented, tested, and repeatable.
  • Actual recovery times and data-loss exposure are measured.
  • Systems without viable recovery paths are visible to leadership.
  • High-priority recovery mitigations are implemented.
  • System-owning teams understand and can maintain their recovery procedures.
  • Recovery capability does not depend entirely on one individual.

Compensation & Benefits

  • Base Compensation: Competitive, based on experience, portfolio strength, and geographic location.
  • Milestone Bonuses: 10% milestone bonus awarded to every team member assisting with the MVP build-out.
  • Annual Performance Bonus: Represents a significant percentage of total compensation.
  • Benefits for Domestic Hires: Health, dental, and vision plans.
  • Annual Paid Offsite: Team retreats in Caribbean, Hawaii, ski destinations, and other exciting locations.

Why Join SunCore Digital?

  • High-impact role in a rapidly scaling digital company.
  • Fully remote team with async flexibility.
  • Direct collaboration with executive leadership and best-in-class marketing/design partners.
  • Opportunity to shape Security processes and mentor junior team members.
  • Competitive compensation with performance upside.
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Location

London, England, United Kingdom

Sign up to applySee more jobs like this