Senior Core Infrastructure Engineer
Oracle · Austin, TX, United States ·
- Seniority
- Senior
- Employment
- Full time
- Category
- Devops
Oracle · Austin, TX, United States ·
Blockchain Association · Remote - United States
S27a · Vancouver, Canada
Thomson Reuters · Toronto, Ontario, Canada
HealthEdge · Hyderabad, IN
Join Oracle NetSuite's Infrastructure Visibility team to build cloud-scale platforms for telemetry, data forensics,
and AI-powered search. Work with a massive data footprint to help engineers investigate complex systems
and improve the reliability of cloud services.
Develop software and automation that make this data accessible and searchable. Own defined components
from design and testing through deployment and support, collaborating with experienced engineers on
broader architecture decisions.
We are looking for experience with:
• Experience developing production software in Python, Java, or another general-purpose language, with
automated testing, Git version control, and code review.
• Experience designing highly available (HA), self-healing systems or services, including failure detection,
retries, failover, and automated recovery.
• Experience with load testing and performance analysis, including interpreting latency, throughput, resource
utilization, and bottlenecks.
• Linux troubleshooting skills covering processes, memory, storage, permissions, and networking.
• Experience developing or integrating REST APIs and working with TLS, certificates, and secure network
connections.
• Experience deploying or operating cloud or containerized services using infrastructure or configuration
automation.
• Clear communication and the ability to collaborate across teams, explain tradeoffs, and carry defined
engineering tasks through implementation and production validation.
Preferred qualifications
Experience with every listed technology is not expected; relevant experience with comparable tools is welcome.
• Experience with telemetry, search, or streaming platforms such as OpenSearch, Elasticsearch, Logstash,
or Kafka.
• Experience with Kubernetes, Helm, GitOps, and cloud infrastructure such as Oracle Cloud Infrastructure or
another major cloud platform.
• Familiarity with Prometheus, Grafana, instrumentation, and actionable alert design.
• Experience with Terraform, Salt, Ansible, or comparable infrastructure and configuration automation tools.
Design, build, and improve highly available, self-healing systems and services using fault detection,
redundancy, automated recovery, and safe handling of partial failures.
• Automate infrastructure, configuration, onboarding, and releases using infrastructure as code and
continuous integration and delivery. Build validation checks, staged rollouts, and recovery procedures.
• Design and execute load tests, analyze throughput, latency, and resource use, identify bottlenecks, and
validate performance and capacity improvements.
• Use telemetry, dashboards, and alerts to investigate system behavior and diagnose production issues. Turn
recurring failures into tested improvements to reliability and operability.
• Develop and integrate services through secure APIs. Review code, document decisions and runbooks, and
partner with application, site reliability, security, and infrastructure teams.
cuculus-gmbh · Leipzig