Machine Learning Evaluation Engineer
Is this the right opportunity for you?
Explore more jobs, compare advertised salaries and see which skills employers want.
Explore careersSearch for a different role
More Data jobs →Skills
Description
The Data, Analytics and Quality (DAQ) team’s core mission is to evaluate and elevate advanced sensing technologies. We collaborate closely with partners in computer vision, video engineering, and other applied ML domains to deliver high quality algorithms that power intelligent features across Apple’s products. Our evaluations and analyses inform these teams and Apple leadership throughout the entire development cycle, from early prototyping to shipping a refined product. We are seeking a Machine Learning Evaluation Engineer who will combine deep technical expertise, creativity, and systems thinking to design and deploy new evaluation strategies for multi-modal foundation models and task-specific CV/ML algorithms.
As a key member of this team, you will lead the benchmarking of state-of-the-art multi-modal models, developing comprehensive evaluation systems that include metric design, large-scale data preparation, and advanced failure analysis automation. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously assess model capabilities and translate findings into actionable improvements. You will leverage modern GenAI tools to improve engineering efficiency and accelerate analysis. You will collaborate with multidisciplinary teams across HW, SW, design, and other applied fields to ensure models meet the high quality bar necessary for exceptional customer experiences.
In this role you'll be responsible for:
Algorithm Evaluation & Benchmarking: Design and build comprehensive evaluation pipelines that scale across large datasets. You will be responsible for both holistic end-to-end system evaluation and granular component-level testing to rigorously measure model capabilities on complex text, image, and video understanding tasks.
Subjective Evaluation: Solve the challenge of developing and calibrating subjective evaluations for generative models by leveraging LLM-as-Judge and human grading techniques.
Deep Failure Analysis: Identify trends in large datasets and dive deep into specific failure cases to find their root causes in models, prompts, or upstream software and algorithm components. Develop failure taxonomies to track over time.
Workflow Automation: Design and develop cloud-based, LLM-powered workflows and dashboards to streamline failure analysis, reporting, and iterative experimentation. This is key to efficiently isolate and triage problems within complex multi-stage algorithm stacks.
Evaluation Set Curation: Refine our strategy for curating high-quality datasets and ground truth, whether through auto-labeling with VLMs, leveraging manual annotation teams, or generating synthetic data.
Data Analysis: assess quality, diversity, representativeness, and coverage gaps across large, complex datasets to ensure evaluation sets reflect real-world usage.
Cross-Functional Collaboration: Partner closely with the core model training teams. You will provide them with actionable, data-driven insights and metrics to guide the next iteration of model training and fine-tuning.
Minimum Qualifications
- BS and a minimum of 3 years relevant industry experience
- 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation.
- Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance. Ability to apply statistical rigor to ensure evaluation metrics are meaningful and reliable.
- Experience analyzing data quality, diversity, representativeness, and coverage gaps in large datasets, and translating findings into data collection and curation strategies.
- Applied experience employing and evaluating generative models, including knowledge of prompt tuning, subjective evaluation of open-ended outputs, and human-in-the-loop methodologies.
- Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for foundation models. Deep experience with custom benchmark creation, automated regression testing, and LLM/VLM-as-a-Judge methodologies.
- Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments. Ability to translate findings into actionable recommendations for model improvement.
- Strong proficiency in Python and building well-designed automated pipelines, including model endpoint integration, multi-step evaluation orchestration, and robust logging for analysis and traceability.
- Experience using AI-powered development and analysis tools to accelerate data exploration, coding, and workflow automation, while maintaining technical rigor and reproducibility.
Preferred Qualifications
- Theoretical and practical understanding of Computer Vision (CV) and Vision-Language Models (VLMs) in order to effectively anticipate failure modes and accurately benchmark their performance.
- Excellent written and verbal communication skills, with the ability to describe complex topics to cross-functional teams with varying levels of technical expertise.
Get similar jobs in United States by email
We'll email you when new jobs similar to this one appear.
Similar jobs
Explore more Data jobs in United States.
Finding similar jobs…