Rodeo
Get started

A&O Shearman

Cloud Operations - Service Reliability Engineer

Belfast
Posted about 17 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

What you will do

The Service Reliability Engineer is accountable for improving the reliability, observability, and operational resilience of cloud-hosted services. The role focuses on monitoring, early issue identification, cloud engineering, and automation, using tools and practices such as Bicep, Azure DevOps, GitHub, and Ansible to support consistent, repeatable, and well-governed service operation.

  • Monitoring, observability, and alerting across cloud infrastructure, platform services, and supported application environments;
  • Cloud engineering and automation, including Infrastructure as Code, deployment pipelines, configuration management, and standards-led delivery; and
  • Promoting service resiliency through proactive issue identification, operational insight, automation, and continuous improvement.

The role involves:

  • Supporting service reliability, observability, and cloud engineering across the following areas:

    • Monitoring, observability, and alerting for cloud-hosted services, including infrastructure health, service availability, performance signals, and operational events - Essential
    • Azure public cloud engineering, including IaaS, PaaS, networking, identity, RBAC, and platform diagnostics - Essential
    • Infrastructure as Code and automation using Bicep, Azure DevOps pipelines, and GitHub-based source control and collaboration - Essential
    • Configuration management and standards automation using Ansible or equivalent tooling - Preferred
    • Experience of using or implementing monitoring solutions using Elastic - Preferred
    • Operational reporting, issue trend analysis, and the development of actionable dashboards to support service improvement - Preferred
    • Experience of working across a broad range of systems, technologies, and internal support teams - Preferred
  • Ensuring that monitoring and operational insight are effectively designed, implemented, and understood so that services can be supported, improved, and made more resilient.

  • Providing subject matter expertise in cloud operations, observability, automation, and reliability engineering practices.

  • Working globally across cloud-hosted services and platform capabilities, independent of location.

  • Supporting the firm's environmental goals and initiatives.

Monitoring, Reliability and Cloud Engineering

  • Works with internal technology teams to improve end-to-end observability for supported services, including:
    • Monitoring coverage for infrastructure, platform services, and application components;
    • Actionable alerting that supports early identification of degradation, failure, or operational risk;
    • Dashboards and reporting that help teams understand service health, trends, and recurring issues;
    • Cloud engineering practices that use Bicep, Azure DevOps, GitHub, and Ansible to deliver consistent and repeatable change; and
    • Operational standards that improve service resilience and reduce manual support effort.
  • Maintains appropriate documentation, including monitoring standards, known issues, operational patterns, troubleshooting guidance, and support handbooks.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Service Delivery

  • Identify, diagnose, and support the resolution of incidents and problems by interpreting monitoring signals, operational telemetry, and service behavior.
  • Work with service-owning teams to improve the quality, relevance, and routing of alerts so that operational issues can be detected and acted upon quickly.
  • Contribute to root cause analysis, problem management, and continuous improvement activity by identifying recurring patterns, gaps in observability, and opportunities for automation.

Build and Implementation

  • Provide specialist guidance to teams adopting cloud engineering patterns, Infrastructure as Code, deployment pipelines, and automated configuration management.
  • Support the implementation of monitoring and automation standards across new and existing services.
  • Ensure that operational documentation, handover materials, and support guidance are created and are suitable for BAU operation.

Risk Management

  • Identify operational, reliability, and supportability risks arising from gaps in monitoring, alerting, automation, or cloud platform standards.
  • Refer to domain experts for guidance on specialized areas such as architecture, security, networking, database platforms, and application design.
  • Participate in recovery, resilience, and operational readiness activities to help prove that services can be supported effectively.

Quality, Methods & Tools

  • Strive for improvements to processes by promoting standardized patterns, automated controls, repeatable engineering practices, and effective use of industry best practice.
  • Advocate for the use of source control, pipeline-based delivery, Infrastructure as Code, and configuration management to improve quality, auditability, and operational reliability.

What you will have

Business Competencies

  • Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis, and service improvement.
  • Technically curious, with a enthusiasm for understanding a broad set of systems, technologies, and operational domains.
  • Ability to interpret monitoring data, identify patterns, and translate operational insight into meaningful improvement activity.
  • Ability to make sound decisions under pressure and support effective incident response.
  • Strong commitment to service reliability, operational resilience, and excellent customer service.
  • Commercial acumen, including an understanding of IT service costs, cloud consumption, and how technology adds value to the business.
  • Ability to promote technical standards, automation, and reliability practices using clear, business-friendly language.
  • Personal credibility; highly self-motivated self-starter who will undertake all activities to the highest professional standards.
  • Excellent communication skills, both orally and written.
  • Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
  • Ability to build effective working relationships across a diverse set of internal teams and influence the adoption of monitoring, automation, and cloud engineering standards.
  • Experience of working in a global environment across international locations with an appreciation of multiple cultures.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Knowledge

  • Practical knowledge of SRE principles, observability, incident response, problem management, and operational resilience.
  • Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking, and automation.
  • Knowledge of Infrastructure as Code, source control, and pipeline-based delivery using tools such as Bicep, Azure DevOps, and GitHub.
  • Knowledge of configuration management and automation tooling such as Ansible.
  • Expected to develop a broad understanding of the technologies, systems, and business working practices used by A&O Shearman.

Experience

  • Minimum 4-5 years' IT experience with at least 2 years' experience in a cloud operations, infrastructure, platform engineering, SRE, or 3rd line support role.
  • Experience of monitoring, alerting, incident investigation, and operational issue identification in a complex technology environment.
  • Experience of using or implementing monitoring using Elastic is desirable.
  • Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub, and Ansible.
  • Experience working with diverse internal teams to improve service supportability, resilience, and operational standards.
  • Experience of working in an ITIL environment.

Qualifications

  • Ideally the candidate should have the following, or equivalent:
    • Minimum "A" level standard education or equivalent;
    • Accreditation in relevant technologies - preferred; and
    • ITIL Foundation - preferred.

Ideal candidate profile

The ideal candidate will be technically curious and motivated by understanding how different systems, platforms, and operational processes fit together. They will enjoy working across a broad technology estate, using monitoring data and engineering insight to identify issues early, improve service resilience, and guide teams towards standardized, automated, and supportable ways of working.

NO AGENCIES PLEASE - A&O Shearman does not accept unsolicited CVs. For further information, please see our UK Recruitment Agency Policy and our commitment to direct sourcing here.

Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Cloud Operations
Service Reliability Engineering
Azure Public Cloud
Infrastructure as Code
Bicep
Azure DevOps
GitHub
Ansible
Monitoring
Observability
Incident Response
Problem Management
Configuration Management
Elastic
ITIL
SRE Principles

Location

Belfast, Northern Ireland, United Kingdom

Sign up to applySee more jobs like this