Sr. Machine Learning Engineer, Speech LLM Evaluation
Apple · Cupertino ·
- Seniority
- Senior
- Category
- ML ai
- Company size
- 1000+
Apple · Cupertino ·
Disney Experiences · Orlando, Florida, United States of America
Amgen · India - Hyderabad
knowtex · San Francisco
Cantina · Europe
Senior ML engineer on Apple's Siri Speech Evaluation team who owns the datasets, metrics, and automated judges used to evaluate speech LLMs (ASR, TTS, real-time conversational models) for accuracy, robustness, and conversational quality before they ship. Core stack is Python, large-scale data pipelines (e.g., Spark), and LLM/human evaluation methods.
Join the team redefining what a deeply personal and integrated assistant can be.
As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS.
Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and real-time conversational AI efforts, partnering directly with the modeling teams. We're growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether they're ready. You'll help define how Apple measures a new class of models that listen, speak, and reason.
This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs.
This role owns the data and metrics foundation for evaluating speech LLMs (e.g., real-time speech understanding and generation models) across accuracy, robustness, and conversational quality. You'll build and curate evaluation datasets that reflect real usage — from personalized named-entity queries to multi-turn fluid conversations — and design the metrics and automated judges that turn model outputs into actionable, trustworthy signal. You'll work closely with modeling, infrastructure, and product partners to make sure every new model is evaluated quickly, consistently, and at the right level of rigor before it reaches customers.
pitchbookdata · Seattle, Washington, United States