SRE
Fulcrum Digital · Ciudad de México, Mexico ·
- Category
- SRE
- Experience
- 25+ years
Fulcrum Digital · Ciudad de México, Mexico ·
StarHub
sailssoftware · Andhra Pradesh, Visakhapatnam, India
SailPoint · Remote (India)
Sherweb · Quebec, Canada
Site Reliability Engineer owning the health, stability, and performance of a production environment supporting a client engagement — defining monitoring/alerting strategies, automating deployments and operations via CI/CD, and reducing incidents and MTTR. Core stack: Linux, Kubernetes, AWS, Jenkins, Splunk/Dynatrace, SQL, with rotational on-call.
Who We Are
Fulcrum Digital is a global
AI-first enterprise transformation company with over 25 years of experience. We
partner with enterprises across financial services, insurance, healthcare,
retail, manufacturing, higher education, and logistics to move from AI experimentation
to scalable business outcomes. Fulcrum Digital works with over 100 global
clients, including Fortune 500 enterprises, combining deep industry expertise
with capabilities in enterprise AI, digital engineering, cloud modernisation,
platform integration, and generative AI.
As a Site Reliability Engineer,
you will own the health, stability, and performance of a production environment
supporting one of our client engagements. You will define how applications are
monitored and supported, drive automation across deployment and operations, and
work closely with development teams to reduce incidents and improve resiliency
over time. This role suits someone with a strong service ownership mindset who
enjoys connecting the dots across a complex technology stack and working with a
global team across multiple time zones.
• Plan, manage, and oversee all aspects of the production
environment.
• Define strategies for application performance
monitoring and optimization in production.
• Design, develop, and standardize monitoring and
alerting mechanisms for supported applications.
• Respond to incidents, improve the platform based on
feedback, and measure the reduction of incidents over time.
• Take a holistic approach to problem solving during
production events, connecting the dots across the full technology stack to
optimize mean time to recover (MTTR).
• Analyze ITSM activities for the platform and provide a
feedback loop to development teams on operational gaps or resiliency concerns.
• Engage in and improve the whole lifecycle of services,
from inception and design through deployment, operation, and refinement.
• Support services before they go live through system
design consulting, capacity planning, and launch reviews.
• Maintain live services by measuring and monitoring
availability, latency, and overall system health.
• Support code deployments into multiple lower
environments, supporting current processes while automating wherever possible.
• Support the application CI/CD pipeline for promoting
software into higher environments through validation and operational gating,
and lead on DevOps automation and best practices.
• Scale systems sustainably through automation, pushing
for changes that improve both reliability and velocity.
• Collaborate with a global team spread across tech hubs
in multiple geographies and time zones, sharing knowledge and explaining
processes and procedures to others.
Requirements
• Strong hands-on experience with Linux.
• Experience with monitoring tools such as Splunk,
Dynatrace, or equivalent.
• Working knowledge of ITIL/ITSM practices.
• Strong troubleshooting skills across complex,
multi-layered platforms.
• Proficiency in SQL and PL/SQL.
• Experience with Jenkins and CI/CD pipelines.
• Scripting experience with Groovy, YAML, and Shell.
• Experience with Git and Bitbucket.
• Hands-on experience with Kubernetes and AWS.
• Proven experience in production support leadership,
including runbook and support model creation.
• Experience defining monitoring and alerting strategies.
• Experience with disaster recovery and resiliency
planning.
• Experience with deployment readiness validation and
operational process design.
• Solid background in root cause analysis and problem
management.
• A service ownership mindset with a focus on continuous
improvement and toil reduction.
• Strong communication skills and the ability to share
knowledge across distributed teams.
• Availability to participate in rotational on-call
duties and occasional off-hours work.
high radius · Hyderabad, Telangana, India