Workflow Automation SRE, ASE Data Services
Apple · Seattle ·
- Category
- SRE
- Experience
- 5+ years
- Company size
- 1000+
Apple · Seattle ·
Apple's Services Engineering organization (ASE) is seeking experienced site reliability engineers to build and maintain the workflow orchestration and automation platform for our Data Services fleet. Data Services SRE operates Cassandra, Redis/Valkey, Kafka, Solr, and our Coordination systems (ZooKeeper, etcd, Parallax) across Apple's data centers worldwide. These systems form the platform upon which iCloud and many other internet services at Apple are built. The operations that keep this fleet healthy (rolling upgrades, host replacements, cluster expansions, backup and restore, release qualification) are long-running, multi-step, failure-prone procedures that today rely on internal workflow systems and manual runbooks. You will work to replace these legacy systems with durable, observable, resumable workflows. Your work will reduce on-call toil, cut incident recovery time, and turn tribal operational knowledge into self-documenting code. In ASE, your work benefits hundreds of millions of users and is critical to the reliability of some of the most visible current and future Apple features.
The ASE Data Services workflow effort builds automation that is safe, reliable, observable, and resumable. This work requires an innovative spirit and an extraordinary degree of care and rigor in engineering. These workflows operate on production databases at massive scale, where a mishandled rolling upgrade or host replacement has real customer impact. You will support the workflow engine layer underlying Apples most critical database systems which power all of Apples internet services. You will design and build workers and workflow definitions, composing operations across existing execution layers: host and pod provisioning, cluster topology discovery, and our monitoring and alerting systems. You will migrate operational logic off home-grown workflow tooling onto a durable execution model with retry, signal, query, and full execution history. This role requires excellent communication, the ability to partner closely with developers, SRE, and platform teams, and a high degree of customer focus when engaging with the internal stakeholders who will run these workflows every day. As a distributed team, the ability to work effectively with colleagues based in other locations is essential; experience in this area is a plus. Prior experience building or operating workflow/orchestration systems, or operating distributed databases and storage systems at scale, is recommended.