Improving search at scale with efficient query experimentation
Most teams judge a search change by relevance labels or one site-wide A/B test. Neither tells you whether a single query got better. This talk lays out how to experiment at the level of the individual intent: which metric to trust, what to randomize on, and how to decide when the data is sparse.
It is the method that now runs unattended as autoLoop: two answers to the same intent, tested live, the winner kept.
Where it lives today: reversing diminishing returns