OpenBox Solutions, Inc
Senior Site Reliability Engineer (SRE) – Cyber Recovery & Backup Engineering
Position Overview
We are seeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, infrastructure resiliency, automation, and reliability engineering.
The ideal candidate will be responsible for designing, implementing, and operating highly available, secure, automated, and resilient recovery capabilities that protect critical enterprise services from operational failures, ransomware, cyberattacks, data loss, and other disruptive events.
This role combines traditional Site Reliability Engineering principles—including automation, observability, reliability engineering, scalability, and resilience—with advanced enterprise backup and cyber recovery technologies.
The engineer will work closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services are continuously recoverable, resilient, and validated against defined RTO/RPO objectives.
Required Education
Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
Equivalent professional experience may be considered in lieu of a degree.
Required Experience
7+ years of experience in Backup Engineering, Infrastructure Engineering, Site Reliability Engineering, or a closely related discipline.
5+ years of experience designing and implementing enterprise backup solutions.
3+ years of experience supporting or engineering cyber recovery architectures.
Proven experience applying SRE principles within enterprise infrastructure environments.
Strong understanding of distributed systems, high availability, resiliency, scalability, and disaster recovery architectures.
Experience engineering solutions for ransomware resilience, cyber recovery, and recovery validation.
Experience automating infrastructure, backup, recovery, monitoring, and operational processes.
Core Technical Skills
Enterprise Backup & Recovery
Hands-on experience with one or more enterprise backup and recovery platforms, including:
Cohesity
Dell PowerProtect Data Manager
Dell Data Domain
Dell Cyber Recovery
Rubrik
Commvault
Veritas NetBackup
Veeam
Experience designing and administering backup solutions across:
Physical servers
Virtual environments
Databases
NAS and Object Storage
Enterprise applications
Kubernetes/OpenShift workloads
Cloud-native workloads
SaaS platforms
Strong understanding of:
Backup architecture and strategy
Backup retention and lifecycle management
Backup replication
Backup performance optimization
Encryption
Recovery objectives
Policy-based backup automation
Enterprise RPO/RTO requirements
Cyber Recovery & Cyber Resiliency
Strong experience with enterprise cyber recovery architectures, including:
Air-gapped backup and recovery vaults
Immutable backups and immutable storage
Clean Rooms
Isolated Recovery Environments (IRE)
Dell Cyber Recovery or equivalent cyber vault technologies
Recovery orchestration
Ransomware recovery
Cyber resilience testing
Recovery point validation
Malware scanning and recovery validation
Secure recovery workflows
Recovery readiness assessment
Cyberattack recovery scenarios
Severe-but-plausible cyber event testing
Cloud & Infrastructure
Experience with one or more major cloud platforms:
Microsoft Azure
Amazon Web Services (AWS)
Google Cloud Platform (GCP)
Knowledge of:
Cloud-native backup
Cross-region recovery
Hybrid cloud resiliency
Enterprise storage platforms
High availability architectures
Distributed systems
Experience supporting:
VMware
Microsoft Hyper-V
Kubernetes
OpenShift
Linux
Windows Server
Active Directory
Automation & Infrastructure as Code
Strong scripting and automation experience using:
Ansible
Terraform
Python
PowerShell
Bash
Experience with:
Infrastructure as Code (IaC)
Recovery as Code
Automated recovery runbooks
Policy-based automation
Recovery validation automation
Compliance evidence automation
Automated reporting
Operational toil reduction
Must Have:
DevOps & CI/CD
Experience with:
GitHub
GitHub Actions
CI/CD pipelines
Automation workflows
Source control
Automated infrastructure and recovery processes
Observability & Monitoring
Experience implementing monitoring, alerting, dashboards, and observability for enterprise infrastructure and recovery platforms.
Required or preferred experience with:
Dynatrace
Grafana
Prometheus
Splunk
ELK Stack
Ability to monitor and measure:
Backup success rates
Replication health
Recovery readiness
Storage utilization
Cyber vault health
Infrastructure dependencies
Platform reliability
Recovery performance
Security & Compliance
Strong understanding of enterprise security and cyber resilience concepts, including:
Zero Trust Architecture
NIST Cybersecurity Framework
CIS Controls
Encryption and key management
Identity and Access Management (IAM)
Multi-Factor Authentication (MFA)
Secure recovery processes
Immutable and air-gapped recovery architectures
Key Responsibilities
Site Reliability Engineering
Engineer and maintain highly available, scalable, resilient enterprise platforms using SRE principles.
Define and measure Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for backup and recovery services.
Establish and manage appropriate error budgets.
Continuously improve platform reliability, scalability, performance, and recoverability.
Identify and eliminate operational toil through automation.
Perform detailed Root Cause Analysis (RCA) and implement permanent corrective actions.
Participate in incident response and major incident recovery activities.
Develop and implement proactive reliability and resilience strategies.
Enterprise Backup Engineering
Design, implement, administer, and optimize enterprise backup and recovery solutions across on-premises, cloud, and SaaS environments.
Engineer backup architectures supporting physical, virtual, database, Kubernetes/OpenShift, cloud-native, NAS/Object Storage, and enterprise application workloads.
Design immutable backup architectures to support ransomware resilience.
Optimize backup performance, retention, replication, encryption, and recovery objectives.
Implement policy-based backup automation and lifecycle management.
Ensure backup and recovery capabilities meet defined enterprise RPO and RTO requirements.
Continuously evaluate backup platform reliability, scalability, capacity, and performance.
Cyber Recovery
Design and implement enterprise cyber recovery architectures, including:
Air-gapped recovery vaults
Clean Rooms
Isolated Recovery Environments (IRE)
Immutable storage
Cyber recovery platforms
Develop secure recovery workflows for cyberattack and ransomware scenarios.
Engineer automated malware scanning and recovery validation processes.
Design and test recovery orchestration for severe-but-plausible cyber events.
Support recovery point validation and promotion into production recovery environments.
Collaborate with Cyber Security teams to develop and strengthen ransomware resilience strategies.
Ensure recovery environments remain isolated, secure, and operationally ready.
Automation & Recovery as Code
Develop Infrastructure as Code (IaC) and Recovery as Code solutions.
Build automated recovery runbooks using Ansible, Terraform, PowerShell, Python, and GitHub Actions.
Automate recovery validation, reporting, compliance evidence generation, and operational workflows.
Eliminate manual recovery processes wherever practical.
Build reusable automation frameworks that improve recovery speed, consistency, and reliability.
Monitoring & Observability
Establish proactive monitoring, alerting, and observability for backup, infrastructure, and cyber recovery platforms.
Develop dashboards providing both operational and executive visibility.
Integrate enterprise observability platforms such as Dynatrace, Grafana, Prometheus, and Splunk.
Monitor backup success, replication health, recovery readiness, storage utilization, cyber vault health, and infrastructure dependencies.
Establish reliability metrics and operational KPIs to identify risks before they become incidents.
Recovery & Resilience Testing
Plan and execute comprehensive cyber recovery and resiliency exercises.
Conduct:
Cyber recovery exercises
Clean Room validation
Air-gap recovery testing
Isolated Recovery Environment (IRE) exercises
Bare Metal Recovery (BMR) testing
Disaster Recovery (DR) testing
Validate application recoverability against defined RTO/RPO objectives.
Test recovery processes against ransomware and other cyberattack scenarios.
Identify recovery gaps and implement corrective actions.
Produce executive-level reporting on recovery readiness, resiliency, and testing outcomes.
Cross-Functional Collaboration
Partner closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams.
Lead cross-functional technical recovery efforts during major incidents and cyber recovery exercises.
Collaborate with application and infrastructure teams to ensure critical workloads are recoverable.
Influence engineering standards, resiliency practices, automation strategies, and operational excellence.
Communicate complex technical recovery and resiliency concepts effectively to both technical teams and executive stakeholders.
Preferred Qualifications
Experience in financial services or another highly regulated industry.
Experience supporting GSIB cyber resiliency programs.
Knowledge of regulatory expectations from organizations such as:
Federal Reserve
Office of the Comptroller of the Currency (OCC)
Federal Financial Institutions Examination Council (FFIEC)
Experience with chaos engineering and resilience testing.
Familiarity with SRE tooling, reliability metrics, SLOs, SLIs, and error budgets.
Experience implementing AI-assisted operations (AIOps) and predictive analytics.
Strong systems-thinking and engineering mindset.
Excellent troubleshooting and Root Cause Analysis (RCA) skills.
Proven ability to lead cross-functional technical recovery initiatives.
Strong communication, presentation, and executive engagement skills.
Ability to influence engineering standards and drive operational excellence.
Demonstrated commitment to continuous improvement through automation and reliability engineering.
Ideal Candidate Profile
The ideal candidate is a senior-level infrastructure/SRE professional who combines:
Enterprise Backup Engineering
Cyber Recovery & Ransomware Resilience
Site Reliability Engineering
Infrastructure & Cloud Engineering
Automation & Infrastructure as Code
Observability & Monitoring
Disaster Recovery
Recovery Orchestration
Security & Cyber Resilience
The candidate should be comfortable operating at both the hands-on engineering level and the architectural level, with the ability to design resilient recovery solutions, automate complex processes, troubleshoot critical infrastructure issues, lead recovery exercises, and communicate recovery readiness to senior leadership.
Key Technology Keywords
SRE | Enterprise Backup | Cyber Recovery | Cyber Resiliency | Ransomware Recovery | Disaster Recovery | Recovery Orchestration | Immutable Backup | Air-Gapped Vaults | Clean Rooms | Isolated Recovery Environments | Cohesity | Dell PowerProtect Data Manager | Dell Data Domain | Dell Cyber Recovery | Rubrik | Commvault | Veritas NetBackup | Veeam | Azure | AWS | GCP | VMware | Hyper-V | Kubernetes | OpenShift | Linux | Windows Server | Active Directory | Ansible | Terraform | Python | PowerShell | Bash | GitHub | GitHub Actions | CI/CD | Dynatrace | Grafana | Prometheus | Splunk | ELK | ServiceNow | Zero Trust | NIST | CIS Controls | IAM | MFA | IaC | Recovery as Code | SLO | SLI | Error Budgets | RCA | RPO | RTO | BMR | AIOps
To apply for this job email your details to bobby@openboxsolutions.com