Site Reliability Engineer (SRE)
Infinity Quest UK · UK ·
- Category
- SRE
This SRE role drives IT operations modernization through observability, automation, and AIOps — building monitoring platforms, AI-driven alerting, and SLO/error-budget practices while leading incident response and mentoring teams. Core tech includes Dynatrace, Datadog, Python, Ansible, AWS, Azure, Docker, and Kubernetes.
Job Description: Key Responsibilities: Drive modernization of IT operations through improved observability and automation. Design and implement observability platforms to monitor system health, performance, and reliability. Develop strategies for AI-driven alerting and proactive anomaly detection to improve MTTD and MTTR. Establish and implement SRE practices, including SLOs, SLIs, and Error Budgets. Develop an AIOps roadmap to improve operational efficiency. Automate repetitive operational tas…