Runware
Senior Site Reliability Engineer

How your CV stacks up
Upload your CV to see how well it fits this job role
?%
About Runware
Runware is building high-performance infrastructure and products to power the world's intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.
About the Role
As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.
What You’ll Do
- Own and improve the reliability, availability and performance of critical production services across the Runware platform
- Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
- Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
- Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
- Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
- Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows
Reasons to use Rodeo
I’m in my final year doing Economics and I don’t know whether to apply for grad schemes now or do a masters first. What do you think?
Honest answer — it depends on where you want to end up. A lot of top grad schemes (Big 4, civil service, banking) don’t need a masters. Let’s look at the ones you’d be competitive for now, and we can decide if a masters actually adds anything.
Also worth knowing: most autumn 2026 applications are open now. Timing matters more than you think.
Start with a chat, not a search bar
Grad scheme, placement, apprenticeship? Not sure what you want yet — that's fine. Your agent talks it through with you and turns "I have no idea" into a shortlist.
Graduate Consultant — 2026 Scheme
Why you're a good match
StrongYour economics background and your summer at a regional bank line up with what PwC looks for on the consulting scheme. Applications close in four weeks.
See breakdownIt searches the market for you
Every day your agent scans the market matching roles against what actually matters to you, not just keywords on a CV.
Why you're a good match
You’ve got the grades and the economics background, and your bank internship is exactly the experience this scheme looks for. Apply soon — deadlines close within the month.
Experience fit
Your summer at the bank plus your econometrics coursework map directly to the day-one responsibilities on this scheme — client modelling, market briefings, and deal support.
Only hits
No noise. No "maybe this fits." Just roles with a clear explanation of why they're right — and where to focus when applying.
Requirements
- Strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
- Strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
- Experience designing and operating observability systems using metrics, logs and distributed tracing
- Understanding SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
- Experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
- Strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation
Bonus
- Experience operating high-throughput or low-latency APIs and distributed systems
- Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
- Experience with RabbitMQ or other distributed messaging and queueing systems
- Experience operating MySQL, Redis, ClickHouse or similar production data systems
- Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
- Experience building automated scaling, capacity management or self-healing systems


Get help with your application
Your very own career expert that helps elevate your application to the next level.
Benefits
- We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time.
- We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.
- Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.
- Generous paid time off – vacation, sick days, public holidays
- Meaningful stock options – share in the upside you create
- Remote-first setup – work from home anywhere we can employ you
- Flexible hours – own your schedule outside core collaboration blocks
- Family leave – paid maternity, paternity, and caregiver time
- Company retreats – twice-yearly gatherings in inspiring locations
“It took my CV and asked me questions relevant to understanding what kind of jobs to suggest for me. Suggestions were almost perfect. Jobs were exactly what I’ve been looking for.”
Jessica, London
Skills
Location