Machine Learning Evaluation Engineer
Apple · Sunnyvale ·
- Category
- ML ai
- Experience
- 3+ years
- Company size
- 1000+
Apple · Sunnyvale ·
OpenAI · Seattle
Survey Monkey · Italy - Remote
weekday-1 · United States
weekday-1 · Bengaluru, Karnataka, India
Designs and deploys evaluation strategies and benchmarking pipelines for multi-modal foundation models and CV/ML algorithms on Apple's DAQ team — covering metric design, large-scale dataset curation, LLM-as-Judge and human grading, and automated failure analysis, then feeding insights back to model training teams.
The Data, Analytics and Quality (DAQ) team’s core mission is to evaluate and elevate advanced sensing technologies. We collaborate closely with partners in computer vision, video engineering, and other applied ML domains to deliver high quality algorithms that power intelligent features across Apple’s products. Our evaluations and analyses inform these teams and Apple leadership throughout the entire development cycle, from early prototyping to shipping a refined product. We are seeking a Machine Learning Evaluation Engineer who will combine deep technical expertise, creativity, and systems thinking to design and deploy new evaluation strategies for multi-modal foundation models and task-specific CV/ML algorithms.
As a key member of this team, you will lead the benchmarking of state-of-the-art multi-modal models, developing comprehensive evaluation systems that include metric design, large-scale data preparation, and advanced failure analysis automation. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously assess model capabilities and translate findings into actionable improvements. You will leverage modern GenAI tools to improve engineering efficiency and accelerate analysis. You will collaborate with multidisciplinary teams across HW, SW, design, and other applied fields to ensure models meet the high quality bar necessary for exceptional customer experiences.
In this role you'll be responsible for:
Algorithm Evaluation & Benchmarking: Design and build comprehensive evaluation pipelines that scale across large datasets. You will be responsible for both holistic end-to-end system evaluation and granular component-level testing to rigorously measure model capabilities on complex text, image, and video understanding tasks.
Subjective Evaluation: Solve the challenge of developing and calibrating subjective evaluations for generative models by leveraging LLM-as-Judge and human grading techniques.
Deep Failure Analysis: Identify trends in large datasets and dive deep into specific failure cases to find their root causes in models, prompts, or upstream software and algorithm components. Develop failure taxonomies to track over time.
Workflow Automation: Design and develop cloud-based, LLM-powered workflows and dashboards to streamline failure analysis, reporting, and iterative experimentation. This is key to efficiently isolate and triage problems within complex multi-stage algorithm stacks.
Evaluation Set Curation: Refine our strategy for curating high-quality datasets and ground truth, whether through auto-labeling with VLMs, leveraging manual annotation teams, or generating synthetic data.
Data Analysis: assess quality, diversity, representativeness, and coverage gaps across large, complex datasets to ensure evaluation sets reflect real-world usage.
Cross-Functional Collaboration: Partner closely with the core model training teams. You will provide them with actionable, data-driven insights and metrics to guide the next iteration of model training and fine-tuning.
Reddit · Remote - United States