Conversational AI evaluation
Machine-learned metrics: models that estimate dialogue quality directly from interaction data, rather than hand-written rules.
For a decade, my work has centered on machine learning for large-scale, real-world problems: conversational AI and speech at Apple's Siri and Amazon's Alexa, and before that warehouse robotics at Amazon.
I'm a machine learning scientist and engineer in conversational AI, speech, and NLP. I focus on the hard problems of alignment and evaluation: knowing whether a system used by millions is genuinely working for the people who rely on it.
At Apple, I work on Siri: the models and quality systems behind how it hears, understands, and responds. That spans on-device speech and wake-word models, multi-turn interaction, and Siri's move into Apple Intelligence and generative AI. Before Apple, I spent several years at Amazon. I co-created IQ-Net to estimate the quality of Alexa conversations at scale, and earlier built performance models and ran large-scale experiments for the robotic systems that move packages through its warehouses.
Across conversational AI, e-commerce, and robotics, the constant has been principled research applied to large, real-world systems: the experiments, models, and machine-learned metrics that let teams ship AI with confidence. Lately my research centers on LLM judges: using language models to evaluate other models, and the statistical pitfalls that undermine them when applied naively.
Machine-learned metrics: models that estimate dialogue quality directly from interaction data, rather than hand-written rules.
Wake-word, endpointing, and real-time speech and language systems that run responsively on-device.
LLM-as-judge panels, and where the common shortcuts for scoring open-ended model outputs quietly fail.
Experimentation and validation for production at scale, from voice assistants to warehouse robotics.
Modeling and evaluation for Siri across successive iOS generations, shipping new voice capabilities to hundreds of millions of users.
Machine-learned evaluation of conversational quality for Alexa, at a scale where manual review is impossible.
Performance modeling of Amazon's first robotic package-sortation system before it scaled.
Earlier I wrote hands-on tutorials on A/B testing, recommender systems, and machine learning from the ground up, now archived here ↗.