Andrew Kornilov
Search practitioner and writer focused on e-commerce search, retrieval, and relevance evaluation. Blogs at frutik.medium.com and maintains the Awesome Search knowledge base. Known for a hands-on, “what actually breaks” style — documenting the hacks and dead ends involved in bending evaluation tooling like Quepid toward vector and image search, and more recently for survey work that takes an evaluation question apart rather than answering it as posed.
Known For
- The four-level framing of judge agreement — label, score, ranking, decision — in Do LLM Judges Actually Agree With Us, arguing that “do LLM judges agree with humans” is a malformed question because the levels do not move together
- A multi-part, hands-on series on adapting Quepid for Vector Search Evaluation
- Building quepid-api-unofficial — an unofficial stateless HTTP API wrapper for Quepid
- Building django-dice — a Python/Django implementation of DICE for e-commerce preference inference, extending it toward implicit behavioural signals and user-inspectable preferences
- Contributing fixes to Quepid
Articles
- Do LLM Judges Actually Agree With Us — landscape survey of LLM-judge agreement evidence, from Voorhees (2000) through 2026 multilingual and product-search results
- Why Setting Up Quepid for Vector Search Evaluation Went Wrong
- Oops, I Did It Again
- How to Evaluate Image Search in Qdrant Using Quepid Part 1
- How to Evaluate Image Search in Qdrant Using Quepid Part 2
Key Concepts
- Levels of Judge Agreement
- LLM as Judge
- Inter-Annotator Agreement
- Vector Search Evaluation
- Quepid
- Judgment Lists
- Search Evaluation
- Multimodal Embeddings
- Hybrid Search
Tools Built
- django-dice — LLM preference inference for Django, following DICE