Senior Site Reliability Engineer
Evri ·
- Work mode
- Remote
- Seniority
- Senior
- Category
- SRE
Evri ·
spectris · Home Working, GB
Zoom · San Jose (CA)
N-iX · Brazil
Broadridge · Manila - 6805 Ayala Ave
Senior Site Reliability Engineer (Team Lead)
Remote | £66,000 + Bonus + Benefits
Build Reliability at Scale. Lead Engineers. Drive Modern Cloud Architecture.
Want to shape the reliability, performance and scalability of platforms that support one of the UK's largest parcel delivery networks?
At Evri, Site Reliability Engineering plays a critical role in how we build, deploy and operate technology. We're looking for a Senior Site Reliability Engineer (Team Lead) to lead and develop a team of six engineers while remaining hands-on across AWS, platform engineering, reliability, automation and cloud-native architecture.
This is a highly influential role requiring both strong leadership and deep technical expertise. We're looking for someone who can drive architectural decisions, champion modern cloud practices, lead complex technical initiatives and help shape the future of reliability engineering across the organisation.
You'll work closely with engineering teams to deliver resilient, scalable and observable platforms that support business-critical services used by millions of customers every year.
What You'll Be Doing
Lead, mentor and develop a team of Site Reliability Engineers
Provide technical leadership across cloud, platform and operational engineering initiatives
Define and drive engineering standards, reliability practices and operational excellence
Lead technical design and architecture discussions to ensure solutions are secure, scalable and maintainable
Improve platform reliability, resiliency, scalability and operational efficiency
Drive platform modernisation initiatives including container adoption, automation and observability improvements
Build and maintain monitoring, logging, alerting and observability solutions that provide meaningful operational insight
Define, measure and improve SLIs, SLOs and key service performance metrics
Drive automation initiatives that reduce manual effort and improve software delivery
Support engineering teams in solving complex availability, performance and scalability challenges
Act as a technical escalation point during major production incidents, leading recovery efforts and driving post-incident improvements
Champion security, resilience and engineering best practices across cloud platforms and services
Manage and evolve shared tooling including CI/CD platforms and automation frameworks
Identify technical debt, challenge existing approaches and drive continuous improvement initiatives
Deliver cost optimisation, platform stability and operational efficiency improvements across the estate
What You'll Bring
Essential Experience
Experience leading, mentoring or managing engineers within an SRE, DevOps, Platform Engineering or Cloud Engineering environment
Deep AWS architecture experience designing, building and operating cloud-native platforms at scale
Strong hands-on experience developing and managing cloud infrastructure using AWS CDK and TypeScript
Strong knowledge of AWS networking and security including VPC design, routing, Security Groups, NACLs, IAM and highly available architectures
Experience designing and supporting modern container platforms including Amazon ECS/Fargate and/or Kubernetes (EKS)
Proven experience owning and leading major production incidents, driving root cause analysis and implementing long-term resiliency improvements
Experience implementing and operating observability platforms including monitoring, logging, tracing and alerting solutions
Strong Infrastructure as Code experience using tools such as CloudFormation, AWS CDK and Ansible
Experience building, supporting and optimising CI/CD pipelines using Jenkins, GitLab, Concourse CI or similar technologies
Strong Linux and/or Windows administration experience
Scripting and automation experience using Python, Bash, Go or similar languages
Excellent troubleshooting and problem-solving skills within complex production environments
Strong stakeholder management and communication skills with the ability to influence technical and non-technical audiences
Demonstrable experience guiding architectural decisions, identifying technical debt and driving platform modernisation initiatives
Desirable Experience
Experience operating distributed systems at enterprise scale
Experience with service mesh, platform engineering or internal developer platforms
Experience working within high-availability, customer-facing environments
FinOps and cloud cost optimisation experience
Why Evri?
At Evri, reliability matters.
You'll join a business where engineering is central to our success and where technology powers millions of parcel deliveries every year. This role offers the opportunity to influence engineering strategy, define reliability standards, lead talented engineers and shape the future of our cloud platforms.
We're looking for someone who doesn't just keep systems running, but actively improves them. Someone who can challenge, innovate and lead from the front.
In return, you'll have the opportunity to make a genuine impact across a large-scale cloud-first environment while continuing to grow your technical and leadership capability.
Let's Deliver It Together
We are Evri.
Where everyone is welcome. Where great engineering thrives. Where reliability is built by design.
LSEG · IND-BLR-Divyasree Technopolis