HCLTech
Site Reliability Engineer (SRE) - Observability & Monitoring Platform

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
About the Role
We are looking for a talented Site Reliability Engineer (SRE) to join a global Platform Services & Engineering team responsible for delivering enterprise-wide monitoring, observability, and notification services.
This role combines production support, platform administration, and reliability engineering. You will help operate and evolve a modern observability ecosystem while partnering with engineering, infrastructure, and support teams to improve platform reliability, scalability, and performance.
This is an excellent opportunity for someone passionate about monitoring, telemetry, automation, and operational excellence, working with state-of-the-art technologies in a highly collaborative global environment.
Technologies & Platforms
- Grafana LGTM Stack
- Open-source observability and telemetry tools
- Enterprise incident and notification platforms
- Linux environments
- CI/CD ecosystems
- Cloud and container technologies
Key Responsibilities
- Drive adoption of observability and monitoring solutions across the organization.
- Administer and support enterprise monitoring, telemetry, and notification platforms.
- Maintain highly available, resilient, and scalable production environments.
- Monitor platform health and proactively respond to alerts before they become incidents.
- Perform incident management, problem management, root cause analysis (RCA), and post-incident reviews.
- Manage production releases, deployments, and change activities with minimal risk.
- Troubleshoot platform issues and support internal users with technical guidance.
- Collaborate with engineering teams to improve operational efficiency and system reliability.
- Design and implement monitoring solutions that enhance visibility across infrastructure and applications.
- Contribute to architectural discussions and help shape the future observability strategy.
- Promote observability best practices and data-driven operational excellence across teams.
- Participate in Agile ceremonies including sprint planning, reviews, and retrospectives.
- Leverage AI-powered tools to improve operational efficiency and automation.
- Mentor junior team members and contribute to knowledge sharing within the team.
- Work closely with globally distributed teams and stakeholders across multiple regions.
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Required Skills & Experience
Mandatory
- Minimum 2 years of experience administering Grafana or other enterprise observability/monitoring platforms.
- At least 2 years of hands-on Linux administration and troubleshooting experience.
- Experience with Python and/or Ansible.
- Strong production support background including:
- Incident Management
- Problem Management
- Change Management
- Release Management
- On-call Support
- User Support & Alert Response
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent communication and stakeholder management skills.
- Basic understanding of cloud technologies and services.
- Familiarity with CI/CD tools such as GitLab, Jenkins, Ansible, Nexus, or similar.
- Ability to work effectively in a collaborative team environment.


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Preferred Qualifications
- Understanding of OpenTelemetry standards and observability concepts.
- Experience with container and orchestration technologies such as Kubernetes, Docker, EKS, or similar platforms.
- Experience supporting medium-to-large-scale production environments.
- Knowledge of AI-assisted operational tools and productivity platforms.
- Familiarity with ITIL principles and service management practices.
- Working knowledge of SQL and relational databases (MySQL, MSSQL, Sybase, etc.).
- Experience with collaboration and project management tools such as Jira and Confluence.
- Exposure to infrastructure technologies including:
- Messaging Middleware (ActiveMQ, Solace, EMS, Tibco)
- Web Servers
- Load Balancers
- Directory Services
- Experience working in globally distributed teams.
What You'll Gain
- Opportunity to work on large-scale, enterprise observability platforms.
- Exposure to cutting-edge monitoring and reliability engineering practices.
- Collaboration with global engineering and operations teams.
- A culture focused on innovation, automation, operational excellence, and continuous improvement.
- Career growth within a fast-evolving SRE and observability landscape.
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location