Nexcess
Reliability Operations Specialist

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
Reliability Operations Specialist
About the Role
We're looking for a Reliability Operations Specialist to help drive operational excellence across incident management, service reliability, observability, and continuous improvement initiatives. This role serves as a central coordinator and subject matter expert for reliability practices, helping teams improve service stability, reduce operational risk, and strengthen operational readiness across the organisation.
The Reliability Operations Specialist partners closely with engineering, infrastructure, security, and operations teams to facilitate:
- Incident response
- Post-incident reviews
- Corrective action tracking
- Visibility into platform health and reliability
This role does not have direct people management responsibilities but plays a critical role in influencing reliability outcomes through:
- Collaboration
- Process ownership
- Data-driven decision making
Location & Employment Details
- Location: Remote
- Employment: Permanent, Full-time
- Pay Range: $85,000 – $100,000 annually (Final compensation based on factors including location, experience, skills, qualifications, and market conditions.)
What You’ll Do
Incident Management & Operational Excellence
- Participate in major incident response activities and serve as an Incident Commander when assigned.
- Coordinate incident response efforts across multiple teams during service-impacting events.
- Facilitate escalation management, stakeholder communications, and status reporting.
- Support the ongoing improvement of:
- Incident management processes
- Procedures
- Operational readiness
- Drive initiatives to reduce:
- Mean Time to Detect (MTTD)
- Mean Time to Resolve (MTTR)
- Identify opportunities to enhance:
- Operational efficiency
- Reliability
- Service delivery
Post-Mortem Management & Continuous Improvement
- Coordinate and facilitate blameless post-mortem reviews following significant incidents.
- Ensure post-mortems are completed:
- Accurately
- Consistently
- Within established timelines.
- Analyse incident trends to identify:
- Recurring issues
- Systemic risks
- Improvement opportunities.
- Maintain accountability for corrective action tracking and closure.
- Partner with stakeholders to:
- Prioritise reliability-focused improvements
- Foster a culture of learning, accountability, and continuous improvement.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Reliability & Observability
- Partner with engineering teams to define, maintain, and mature:
- Service Level Indicators (SLIs)
- Service Level Objectives (SLOs)
- Support the development and evolution of:
- Observability practices
- Monitoring
- Alerting
- Dashboards
- Telemetry standards
- Observability practices
- Analyse reliability metrics and operational performance data to identify opportunities for improvement.
- Recommend and track initiatives that enhance:
- Platform stability
- Resiliency
- Service performance
- Help establish operational best practices for scalable and reliable platform operations.
Reporting & Stakeholder Communication
- Develop reliability reporting for:
- Engineering leadership
- Executive stakeholders.
- Maintain:
- Incident communication standards
- Stakeholder notification processes.
- Provide regular reporting, including:
- Incident performance
- Corrective actions
- Reliability trends
- Service health.
- Translate technical reliability metrics into:
- Actionable business insights
- Recommendations.
- Present findings and recommendations to:
- Technical and non-technical audiences.
What You’ll Bring
Essential Requirements
- 3+ years of experience in one of the following:
- Product Operations
- Platform Operations
- Technical Customer Support
- Incident Coordination
- Site Operations
- IT Service Management (ITSM)
- Related operational roles.
- Experience participating in or coordinating major incident response activities.
- Knowledge of:
- Incident management
- Root cause analysis
- Problem management
- Post-mortem methodologies.
- Experience working with:
- Monitoring tools
- Alerting tools
- Observability tools
- Operational reporting tools.
- Strong:
- Analytical skills
- Organizational skills
- Exceptional attention to detail.
- Excellent:
- Written communication skills
- Verbal communication skills.
- Ability to:
- Work effectively across multiple teams
- Influence outcomes without direct authority.
- Strong problem-solving skills with the ability to:
- Remain calm and organised during high-pressure situations.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Preferred Qualifications
- Experience with:
- Service Level Objectives (SLOs)
- Service Level Indicators (SLIs)
- Reliability metrics.
- Familiarity with:
- Linux systems
- Cloud infrastructure
- Networking concepts
- Hosting platforms
- Distributed systems.
- Knowledge of:
- ITIL
- Operational excellence frameworks
- Site Reliability Engineering (SRE) principles.
- Experience supporting:
- High-availability SaaS
- Hosting environments
- Cloud infrastructure environments.
- Experience creating:
- Executive-level operational reports
- Dashboards
- Presentations.
- Experience using:
- Observability platforms
- Incident management platforms.
What We Offer
- Comprehensive benefits package
- Traditional and Roth 401(k) with company matching
- A collaborative, team-oriented culture
- Consistent and predictable work hours
- Engaging, varied work that keeps each day different
- Opportunities to:
- Contribute ideas
- Influence how work gets done
Disclaimer
This job description is only a summary of the typical functions of the position. It is not intended to be an exhaustive or comprehensive list of all job responsibilities, tasks, or duties. Additional duties and tasks may be assigned as part of the job function. Nexcess reserves the right to modify, interpret, or apply this job description in a way that best supports the ** organisational needs**.
The job description in no way creates or implies an employment contract. Employment remains “at will”.
Equal Employment Opportunity Policy
Nexcess is committed to offering equal employment opportunity without regard to age, color, disability, gender, gender identity, genetic information, marital status, military status, national origin, race, religion, sexual orientation, veteran status, or any other legally protected characteristic.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location