Rodeo
ResourcesPartnersSign in

Herbert Smith Freehills Kramer

Performance & Observability Engineer

City of London
Posted about 23 hours ago
Sign up to applySee more jobs like this

How your CV stacks up

1Upload CV
2Analyse CV
3Improve CV

Upload your CV to see how well it fits this job role

?%

Role Summary

The Performance & Observability Engineer plays a critical role in ensuring system reliability, scalability, and visibility across the entire technology stack. This role focuses on transitioning from traditional monitoring to full observability, enabling deep performance insights, real-time issue detection, and proactive optimisation. The ideal candidate will have expertise in performance engineering, distributed tracing, logging, metrics, and automation, driving improvements across cloud, infrastructure, and application layers.

Primary Responsibilities

Performance Engineering & Optimisation

  • Analyse application, database, and infrastructure performance to identify bottlenecks and inefficiencies.
  • Develop performance benchmarks and SLIs to measure service responsiveness and stability.
  • Collaborate with SRE and DevOps teams to optimise CI/CD pipelines for performance improvements.
  • Collaborate with wider IT teams to implement caching strategies, query optimisation, and autoscaling to enhance system efficiency.

Observability Platform Development & Implementation

  • Design and implement end-to-end observability frameworks covering metrics, logs, traces, and events.
  • Instrument services using existing tools (e.g.: Nexthink) to improve visibility.
  • Enable distributed tracing across microservices to enhance root cause analysis and performance debugging.
  • Standardise logging and telemetry collection across infrastructure, applications, and cloud services.
  • Define best practices and consistent approach across development teams to improve monitoring consistency.

Maturing from Monitoring to Full Observability

  • Transition from basic alerting to proactive insights, leveraging AI-driven anomaly detection.
  • Ensure comprehensive observability across frontend, backend, databases, cloud infrastructure, and networking.
  • Implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to track system health.
  • Automate root cause analysis and incident detection through advanced monitoring techniques.

Reasons to use Rodeo

I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?

Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.

Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.

Start with a chat, not a search bar

Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.

P

Graduate Consultant — 2026 Scheme

PwC·London, UK
£35,000/yr

Why you're a good match

Strong

Your economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.

See breakdown
Save jobNot relevant
View details

It searches the market for you

Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.

Why you're a good match

You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.

See breakdown
Strong

Experience fit

Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.

See breakdown
Strong

Only hits

No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.

Incident Response & Reliability Engineering

  • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through improved observability.
  • Integrate monitoring and alerting tools with incident response platforms (i.e.: ServiceNow).
  • Develop self-healing and auto-remediation mechanisms to reduce operational toil.
  • Improve alerting strategies by reducing false positives and improving signal-to-noise ratio.

Governance, Compliance & Reporting

  • Define observability best practices and governance models to ensure adoption across teams.
  • Ensure log retention, security, and compliance with standards (e.g., GDPR, SOC 2, PCI DSS).
  • Develop executive dashboards and reporting frameworks to showcase reliability and performance trends.

Key Performance Indicators

Maturity of Observability Capabilities

  • % of Services with Full Observability Coverage - Ensure visibility across the entire stack.
  • Instrumentation Completeness (%) - Track the number of services fully instrumented with logs, metrics, and traces.
  • Service-Level Indicator (SLI) Coverage - Ensure key performance indicators are defined and tracked.
  • Reduction in Blind Spots (%) - Improve monitoring coverage across all components.

Performance & Reliability Metrics

  • Application Response Time (P99, P95, P50 Latency) - Improve service performance.
  • System Throughput & Load Handling (%) - Increase service efficiency and scalability.
  • Reduction in Performance Bottlenecks (%) - Optimise infrastructure and application layers.
  • Successful Load Test Pass Rate (%) - Ensure applications meet expected performance benchmarks.

Incident Management & Operational Efficiency

  • Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) - Reduce system downtime and recovery time.
  • Reduction in Noisy Alerts (%) - Improve alert relevance and reduce false positives.
  • Proactive Issue Detection (%) - Increase the percentage of incidents identified before user impact.
  • Error Budget Utilization (%) - Ensure system reliability is balanced with innovation velocity.

Get help with your application

Your very own career expert that helps elevate your application to the next level.

Get help applying for this job

Automation & AI-Driven Observability (AIOps)

  • % of Issues Resolved via Automated Remediation - Reduce manual intervention in incident response.
  • Reduction in On-Call Burden (%) - Minimize alerts requiring human intervention.
  • Anomaly Detection Accuracy (%) - Improve proactive detection of performance issues.

Business & Customer Impact

  • Customer Experience Metrics (Latency, Errors, Uptime) - Ensure observability drives tangible business improvements.
  • Downtime Reduction (%) - Improve service availability and reliability.
  • Cost Optimisation from Performance & Observability (%) - Reduce operational inefficiencies and cloud expenses.

Qualifications, Skills and Experience

Technical Skills

  • Experience in using and maintaining Observability & APM Tools - Grafana experience is essential.
  • Experience of using KQL.
  • Performance Testing & Load Testing.
  • Cloud & Infrastructure Monitoring - AWS CloudWatch, Azure Monitor, GCP Operations Suite, Kubernetes Observability.
  • Knowledge of Log Aggregation & Analysis - e.g.: Grafana Loki, Splunk etc.
  • Automation & Scripting - PowerAutomate, Terraform, PowerShell.
  • Incident Response & ITSM - ServiceNow.
  • Sound understanding of firm's applications, systems and tools across technology stack.
  • DevOps experience is desirable.
  • Nexthink experience is desirable.

Soft Skills & Collaboration

  • Strong problem-solving and root cause analysis skills.
  • Ability to translate observability insights into business impact for stakeholders.
  • A continuous improvement mindset, focused on reducing toil and improving efficiency.
  • Experience working in a DevOps, SRE, or Platform Engineering environment.
Trusted by 25,000+ job seekers

“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”

Jessica, London

Get help applying for this job

Skills

Performance Engineering
Observability
Distributed Tracing
Grafana
KQL
AWS CloudWatch
Azure Monitor
GCP Operations Suite
Kubernetes
Grafana Loki
Splunk
Terraform
PowerShell
ServiceNow
Nexthink
DevOps

Location

City of London, England, United Kingdom

Sign up to applySee more jobs like this