Site Reliability Engineer, AI Platform
XenonStack · Mohali, India ·
- Category
- SRE
- Experience
- 3+ years
XenonStack · Mohali, India ·
XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.
We build enterprise-grade platforms across the agentic stack:
Our mission is to accelerate the world’s transition to AI + Human Intelligence by making agentic systems reliable, responsible, and enterprise-ready.
We are seeking an Site Reliability Engineer, AI Platform to design and implement end-to-end observability frameworks for AI-native and multi-agent systems.
This role sits at the heart of AgentOps and Reliability Engineering — ensuring that agents, pipelines, and infrastructure are monitored, measurable, and continuously optimized.
If you thrive on metrics, monitoring, and making complex systems transparent and reliable, this role offers a chance to define observability for the next generation of enterprise AI.
Observability Frameworks
Design and implement observability pipelines covering metrics, logs, traces, and cost telemetry for agentic systems.
Build dashboards and alerting systems to monitor reliability, performance, and drift in real-time.
Agentic AI Monitoring
Track LLM usage, context windows, token allocation, and multi-agent interactions.
Build monitoring hooks into LangChain, LangGraph, MCP, and RAG pipelines.
Reliability & Performance
Define and monitor SLOs, SLIs, and SLAs for agentic workflows and inference infrastructure.
Conduct root cause analysis of agent failures, latency issues, and cost spikes.
Automation & Tooling
Integrate observability into CI/CD and AgentOps pipelines.
Develop custom plugins/scripts to extend observability for LLMs, agents, and data pipelines.
Collaboration & Reporting
Work with AgentOps, DevOps, and Data Engineering teams to ensure system-wide observability.
Provide executive-level reporting on reliability, efficiency, and adoption metrics.
Continuous Improvement
Implement feedback loops to improve agent performance and reduce downtime.
Stay updated with state-of-the-art observability and AI monitoring frameworks.
Must-Have
3–6 years of experience in SRE, DevOps, or Site Reliability Engineer, AI Platforming.
Strong knowledge of observability tools (Prometheus, Grafana, ELK, OpenTelemetry, Jaeger).
Experience with cloud-native infrastructure (AWS, GCP, Azure) and Kubernetes monitoring.
Proficiency in Python, Go, or Bash for scripting and automation.
Understanding of AI/LLM pipelines, RAG systems, and vector databases.
Hands-on with CI/CD pipelines and monitoring-as-code.
Good-to-Have
Experience with AgentOps tools (LangSmith, PromptLayer, Arize AI, Weights & Biases).
Exposure to AI-specific observability (token usage, model latency, hallucination tracking).
Knowledge of Responsible AI monitoring frameworks.
Background in BFSI, GRC, SOC, or other regulated industries.
At XenonStack, we believe in shaping the future of intelligent systems. We foster a culture of cultivation built on bold, human-centric leadership principles, where deep work, simplicity, and adoption define everything we do.
Our Cultural Values
Agency – Be self-directed and proactive.
Taste – Sweat the details and build with precision.
Ownership – Take responsibility for outcomes.
Mastery – Commit to continuous learning and growth.
Impatience – Move fast and embrace progress.
Customer Obsession – Always put the customer first.
Our Product Philosophy
Obsessed with Adoption – Making observability and trust an integral part of enterprise AI.
Obsessed with Simplicity – Turning complex monitoring into seamless, actionable insights.
Be part of our mission to accelerate the world’s transition to AI + Human Intelligence — by making agentic AI systems transparent, observable, and reliable at scale.