Rodeo
Get started

Mphasis

Site Reliability Engineer

London
Posted about 17 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Job Description

As an SRE, you'll collaborate closely with Application Development and Operations teams to build and maintain scalable systems. Your core focus will be to automate processes and ensure the highest levels of service reliability, specifically by reducing manual effort (TOIL). You'll bring a strong passion for continually improving the reliability, availability, and performance of our services.

Primary Skill

  • Experience with cloud platforms Primarily in AWS Cloud (e.g., AWS, GCP, Azure) and Container Orchestration (e.g., Kubernetes, Docker).
  • Proficiency in Monitoring and Logging Tools: Datadog, Splunk, Dynatrace, AppDynamics, Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana), Cloude Watch, Gremlin, Thousand Eyes.
  • Terraform, Jenkins, GitLab CI, PostgreSQL, Redis, Kong API.
  • Infrastructure skills, Networking and Security Skills, AWS (Atlas), ECS Based internal tooling
  • Lucidchart, PlantUML

Secondary Skill

  • SNOW, Jira, Shell Script, Linux, bitbucket, Akamai, DevOps

Expectations

  • 6+ years of experience with SRE principals and tools (AWS, Monitoring tool, DevOps, CI/CD etc.) should have worked with Toil identification
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or similar role.
  • Achieve and Maintain the Define SLI, SLO, SLA with business/operations/Engineering team.
  • Monitoring, logging, event detection, Alerting and Error budget (99.9, 99.99, 99.999%) for Cloud or Distributed platforms software, Operations & Business.
  • Good understanding of programming skills in languages such as Python, Go, Java, or Ruby.
  • Experience with cloud platforms Primarily in AWS Cloud (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes, Docker).
  • Skilled in managing configuration, deployments, observability, handling and resolving incidents, including root cause analysis, managing and operating complex systems for scalability, availability and performance.
  • Experience with CI/CD pipelines and DevOps tools.
  • Good Knowledge of Source Code Repository Tools
  • Good Knowledge of networking and Security concepts.
  • Proficiency in Monitoring, Logging & Traceability Tools
  • Good Understanding of Database Systems – Aurora PostgreSQL, REDIS
  • Infrastructure skills: Terraform for infrastructure as code, Skilled in the understanding of use, core cloud application infrastructure services including identity platforms, networking, storage, databases, containers, and serverless
  • Efficiency in creating Dashboard for Infra / APM / E2E workflows.
  • ITIL – Incident/ Change, Proficient in Problem management and Jira – Blameless postmortem, Root Cause Analysis findings, applying permanent fixes, Documentation as Runbooks for lesson learn
  • Hands on experience on technical operations application support and stability, reliability and resiliency
  • Proficient in communication and collaboration skills to work effectively with development and operations teams.
  • Willingness to learn new tools, technologies and adapt to changing environments.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Nature of the Job

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job
  • SRE end-to-end Support & Monitoring as an Engineer and Architecture.
  • Design, implement to maintain scalable and reliable systems for services.
  • Setup and Monitor system performance and reliability, identifying and resolving issues proactively.
  • Assess current state and mature the SRE function.
  • Collaborate with development teams to improve system architecture and design for reliability.
  • Participate and own the issues and incident response efforts to resolve production issues. Resolve incidents and provide Incident Resolution with root cause analysis (RCA).
  • Conduct post-mortem analyses of incidents to identify root causes and implement preventive measures.
  • Create and maintain documentation as a Runbook for systems, processes, and procedures.
  • Advocate for best practices in software development, system design, and operational excellence with Tool Implementation.
  • Create and maintain monitoring technologies and processes that improve the visibility of our applications' performance and business metrics and keep operational workload in-check.
  • Establish and ensure the repeatability, traceability, and transparency of application components.
  • Monitoring and optimization of system performance and resource usage, identify and address bottlenecks, and implement best practices for performance tuning.
  • Troubleshoot and resolve critical issues across multiple layers, including CDN, Load balancers, storage, OS, network, K8S, virtualization, and application/DB stack.
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

AWS
Kubernetes
Docker
Terraform
Python
Go
Java
Ruby
Datadog
Splunk
Prometheus
Grafana
CI/CD
PostgreSQL
Redis
Linux

Location

London, England, United Kingdom

Sign up to applySee more jobs like this