Guneet Singh Kohli Principal Machine Learning Scientist & Engineer · Siri @ Apple

I build and evaluate the AI behind everyday conversations.

For a decade, my work has centered on machine learning for large-scale, real-world problems: conversational AI and speech at Apple's Siri and Amazon's Alexa, and before that warehouse robotics at Amazon.

Now: Siri modeling & evaluation, Apple Focus: Conversational AI · Speech & NLP · Evaluation Based in Seattle
Guneet Singh Kohli, headshot

Aboutwho / what

I'm a machine learning scientist and engineer in conversational AI, speech, and NLP. I focus on the hard problems of alignment and evaluation: knowing whether a system used by millions is genuinely working for the people who rely on it.

At Apple, I work on Siri: the models and quality systems behind how it hears, understands, and responds. That spans on-device speech and wake-word models, multi-turn interaction, and Siri's move into Apple Intelligence and generative AI. Before Apple, I spent several years at Amazon. I co-created IQ-Net to estimate the quality of Alexa conversations at scale, and earlier built performance models and ran large-scale experiments for the robotic systems that move packages through its warehouses.

Across conversational AI, e-commerce, and robotics, the constant has been principled research applied to large, real-world systems: the experiments, models, and machine-learned metrics that let teams ship AI with confidence. Lately my research centers on LLM judges: using language models to evaluate other models, and the statistical pitfalls that undermine them when applied naively.

Focuswhat I work on

01

Conversational AI evaluation

Machine-learned metrics: models that estimate dialogue quality directly from interaction data, rather than hand-written rules.

02

Speech & on-device voice

Wake-word, endpointing, and real-time speech and language systems that run responsively on-device.

03

Evaluating generative AI

LLM-as-judge panels, and where the common shortcuts for scoring open-ended model outputs quietly fail.

04

ML for large systems

Experimentation and validation for production at scale, from voice assistants to warehouse robotics.

Experiencetap to expand

Modeling and evaluation for Siri across successive iOS generations, shipping new voice capabilities to hundreds of millions of users.

  • "Siri" invocation models: moving from "Hey Siri" to just "Siri" without raising false activations. announced ↗
  • Back-to-back requests: letting people make follow-ups without re-triggering the assistant. announced ↗
  • Siri + generative AI: modeling and evaluation behind Siri's integration of large language models, including ChatGPT. announced ↗
  • Adaptive on-device speech endpointing: detecting in real time when a person starts and stops speaking, across accents and speech differences.

Machine-learned evaluation of conversational quality for Alexa, at a scale where manual review is impossible.

  • IQ-Net: a neural architecture for interaction-level dialogue quality, learned directly from raw dialogue content and system metadata rather than hand-engineered NLP features. Flexible enough to swap in different underlying models. Published at KDD Converse 2020. paper ↗
  • End-to-end, semi-supervised learning to detect unsatisfactory Alexa interactions.
  • Subject-matter expert for the AWS Certified Machine Learning exam.
  • Conducted 60+ machine learning and data scientist interviews at Amazon.

Performance modeling of Amazon's first robotic package-sortation system before it scaled.

  • Performance modeling of Amazon's first robotic package-sortation system; identified defects and bottlenecks limiting throughput.
  • A/B testing at production scale: designed and analyzed experiments evaluating new features and robotic technologies across Amazon warehouses.