Software Development Engineer - MultiCloud
Apple · Austin ·
- Category
- Software engineering
- Company size
- 1000+
Apple · Austin ·
Upstart · United States | Remote
thundercompute · San Francisco
JP Morgan Chase · BOURNEMOUTH, DORSET, United Kingdom
Recruiting From Scratch · Nob Hill, San Francisco
Engineer on Apple's MultiCloud team in Austin building and operating reliable, secure cloud platforms: writing services and automation in Python/Go, running Kubernetes, improving observability and reliability, and supporting AI/ML and GPU-based workloads.
We're looking for a motivated Software Engineer to join our team and help build and operate reliable, secure, and scalable cloud platforms and services. In this role, you'll develop software and automation that support production systems, Kubernetes platforms, cloud infrastructure, observability, analytics, and emerging AI/ML workloads.
You'll combine software engineering skills with reliability engineering principles to improve how our platforms are built, deployed, monitored, and operated. You'll work alongside experienced engineers, architects, SREs, and AI/ML teams to solve technical challenges and continuously improve our systems.
The ideal candidate has a strong software engineering foundation, enjoys solving problems through code and automation, and is interested in cloud technologies, Kubernetes, reliability engineering, analytics, and AI.
Software Development: Design, develop, test, and maintain software, services, APIs, tools, and automation using languages such as Python and Go.
Cloud & Platform Engineering: Build and improve cloud-native services and platform capabilities across production and non-production environments.
Kubernetes: Deploy, operate, and troubleshoot applications and services running on Kubernetes. Develop automation that simplifies deployment and platform operations.
Reliability: Apply reliability engineering practices to improve system availability, performance, scalability, and operational efficiency.
Automation: Identify repetitive operational activities and develop software and automation to reduce manual effort and operational toil.
Observability: Use logs, metrics, traces, dashboards, and alerts to understand system behavior, troubleshoot issues, and identify opportunities for improvement.
Analytics: Analyze operational and application data to identify trends, anomalies, recurring issues, and performance bottlenecks. Develop dashboards and reporting that provide actionable insights.
AI/ML Infrastructure: Support infrastructure and platform capabilities for AI/ML training, LLM inference, and GPU-based workloads. Develop automation to simplify deployment and operation of these environments.
AI-Assisted Engineering: Explore and apply LLMs and AI technologies to improve software development, troubleshooting, analytics, automation, and operational workflows. Scale & Resilience: Participate in capacity planning, performance testing, scale testing, and disaster recovery exercises.
Continuous Improvement: Identify opportunities to improve platform reliability, developer experience, automation, and operational processes.
Documentation: Create and maintain technical documentation, operational procedures, troubleshooting guides, and runbooks.
Collaboration: Work closely with software engineering, platform, SRE, QA, AI/ML, security, architecture, and program management teams.
Cloudflare · Hybrid