Site Reliability Engineer part of Oracle Analytics Service Excellence (OASE) SRE team partnering with Oracle Analytics development teams to improve the reliability, availability, performance, operational support and maturity of Oracle Analytics Cloud services.
The ideal candidate is a hands-on Site Reliability Engineer with strong analytical skills and the ability to read, understand, investigate, and safely troubleshoot existing application code. This role requires diagnosing complex service issues across distributed systems, automation, infrastructure, and enterprise applications—using logs, telemetry, code investigation, and operational data to drive issues from detection through mitigation and prevention.
OASE develops tools, technologies, processes, and data-driven operating practices that improve service uptime, reduce time to mitigation, and enable scalable cloud operations. The team builds and operates internal services, automation, dashboards, and reporting capabilities that support Oracle Analytics customers, engineering teams, partners, and business growth. The ideal candidate enjoys working in an agile, customer-focused environment. This role is centered on improving uptime through proactive monitoring, incident response, code-level troubleshooting, automation, capacity analysis, operational tooling, patching, remediation, and continuous service improvement.
Responsibilities
Required Qualifications
- Working knowledge of Java, with the ability to read, understand, and troubleshoot existing application code. Experience with other object-oriented programming languages, such as C#, is also valuable.
- Experience with cloud-native applications, containers, Kubernetes, microservices, and independently scalable services.
- Experience developing, operating, or supporting cloud services and large-scale distributed applications in production.
- Demonstrated ability to troubleshoot complex technical issues methodically, including investigation of existing applications and code.
- Linux/Unix system-administration experience, including troubleshooting processes, memory, CPU, filesystem, and network issues.
- Experience supporting cloud infrastructure, networking, applications, services, tools, and operational processes.
- Strong understanding of networking and TCP/IP fundamentals, including DNS, HTTP/HTTPS, TLS, load balancing, and service connectivity.
- BS or MS in Computer Science, Engineering, or equivalent practical experience.
- Experience creating and maintaining technical documentation, runbooks, knowledge articles, and operational guides.
- Experience working in agile development and operational environments.
- Strong written and verbal communication skills, including the ability to work effectively with remote global teams.
Ability to work independently, manage competing priorities, and participate in on-call, after-hours maintenance, and weekend support as needed.
Preferred Qualifications
- Programming and scripting experience with Python, Bash, JavaScript/Node.js, or similar languages, particularly for automation, monitoring, and troubleshooting.
- Experience with OCI, AWS, Azure, or GCP compute, storage, networking, monitoring, and operational tooling.
- Ability to read, understand, troubleshoot, and safely modify existing enterprise application code.
- Experience with Oracle Analytics Cloud, Oracle Analytics Server, OBIS, BI Publisher, Oracle Database, Autonomous Database, MySQL, SQL Server, or NoSQL technologies.
- Experience with REST APIs, service integrations, and automation workflows.
- Familiarity with AI-assisted development tools, such as Codex and Claude Code, for software development, automation, investigation, and documentation.
- Experience using Jira and Confluence for incident management, issue tracking, operational documentation, and collaboration.
- Two to four years of experience operating large-scale, customer-facing web applications or cloud services.