Awesome Search — Information Retrieval & Search Knowledge Graph

Hello, I am Andrew.

I’ve been building e-commerce search applications for 15+ years. Over that time, I’ve collected and connected ideas from publications, conference talks, books, research papers, blog posts, and practitioners across the information retrieval ecosystem.

This knowledge graph maps many of the resources that have influenced my thinking, organized by topic and interconnected through shared concepts. Because search is inherently multidisciplinary, many resources are linked to multiple areas of the graph, reflecting how ideas from ranking, relevance, user behavior, machine learning, evaluation, and system design often overlap.

⭐ Star us on GitHub — it helps!

Semantic knowledge graph built from the Awesome Search curated list. Contains article notes (for paywalled articles, only summaries and key concepts are included), concept notes, topic notes, people notes, case study notes, and company notes, all interconnected through wikilinks.

Maps of Content (Entry Points)

DomainMOC
Agentic Search & EmbeddingsMOC - Agentic Search and Embeddings
Search Quality & Query UnderstandingMOC - Search Quality Assurance and Query Understanding
Ranking & RetrievalMOC - Ranking and Retrieval
Search UX & DiscoveryMOC - Search UX and Discovery
Case StudiesMOC - Case Studies
Architecture & Search TeamMOC - Architecture and Search Team

Core Concepts by Domain

Retrieval

BM25 · Dense Vector Retrieval · Brute-Force Vector Search · Sparse Vector Retrieval · Learned Sparse Retrieval · SPLADE · miniCOIL · Hybrid Search · Reciprocal Rank Fusion · Relative Score Fusion · Score Normalization · Semantic Boosting · Semantic Search · Zero-Shot Retrieval · Out-of-Vocabulary · SIRA

Embeddings

Bi-Encoder · Cross-Encoder · Dense Passage Retriever · ColBERT · Late Interaction · MUVERA · Pooling · Token Pooling · Matryoshka Embeddings · SPLADE · ELSER · Task-Aware Embeddings · Hypothetical Document Embeddings · Dimensionality Reduction · PCA · t-SNE · UMAP · Vector Quantization · Scalar Quantization · Binary Quantization · TurboQuant · RaBitQ · ASH · ITQ

Ranking

Learning to Rank · Personalization · Position Bias · Impression Bias · Ranking Signal Selection · Hashing Trick · Diversity Metrics · Retrieval Pipeline · Results Boosting · Results Merchandising · Signal Downboosting

Model Training

Embedding Fine-tuning · Contrastive Learning · Knowledge Distillation · Hard Negative Mining · Synthetic Query Generation · Consistency Filtering · PROMPTAGATOR · FLAN-T5

Evaluation

NDCG · MRR · MAP · Precision and Recall · UDCG · Search Evaluation · Judgment Lists · Vector Search Evaluation · LLM as Judge · Staged Judging · Statistical Significance in Search Evaluation · Session-Based Evaluation · Out-of-Time Validation · Interleaving · Isolated Feedback Loops · Click Signals · Pointwise Relevance Evaluation · Pairwise Relevance Evaluation · Listwise Relevance Evaluation

Query Understanding

Query Understanding · Query Types · Search Intent · Query Segmentation · Collocations

Lexical Query Operations

Tokenization · Spelling Correction · Synonyms · Stopwords · Autocomplete · Query Expansion · Query Relaxation

Search UX & Discovery

Search UX · Faceted Search · Search Scopes · Federated Search · Zero Results · Presentation Bias · Results Merchandising · Search Result Diversity · MMR

Architecture & RAG

Search Architecture · Sharding · Compute-Storage Disaggregation · Knowledge Graph Search · RAG · Agentic Search · Search-R1 · Reinforcement Learning for Search · Vector Filtering · Text Chunking · Clean Context

Topics

Practice-oriented guides — how to DO or deal with something in search.

Search Quality Assurance · Quepid Beyond Supported Engines · A-B Testing for Search · Duality in Measuring Search · NDCG Variants · Managing a Search Team · Understaffed Search Team · Hiring for Search · Economics of Search · E-commerce Search · Two-Sided Marketplace Ranking · Autocomplete and Autosuggest · Search Result Diversity · Synonyms and Vocabulary Management · Query Understanding in Practice · Query Classification · Multilingual Search · Relevance Program Setup · Personalization in Search · Conversational and Agentic Search · Spelling Correction in Search · Vector Search Tradeoffs · Dimensionality Reduction vs Quantization · PCA vs t-SNE for Retrieval · Elasticsearch vs OpenSearch · Federated vs Unified Search · Late Interaction in Elasticsearch · Late Interaction in OpenSearch · Late Interaction in Qdrant · Late Interaction in Vespa · Extreme Search Systems · Multi-Tenancy in Search · Migration between Search Engines · Embedding Models Compared · Model Selection and Fine-Tuning Evaluation · Retrieval Benchmarks and Leaderboards

Tools

Quepid · Search Relevance Workbench · Elasticsearch Relevance Studio · User Behavior Insights · Querqy · Elasticsearch · OpenSearch · Solr · Qdrant Vector DB · Weaviate Vector DB · turbopuffer Search DB · FAISS · ann-benchmarks

Companies

Technology Providers Elastic · Vespa · Meta · Cohere · OpenSource Connections · Algolia · Weaviate · searchHub · Empathy · Sease · MongoDB · Voyage AI · Qdrant · Hornet · Amazon Web Services · turbopuffer

End Users Uber · Airbnb · Booking.com · Zalando · Slack · Canva · Netflix · Twitter · Etsy · Skyscanner · Grubhub · Spotify · Carousell · Vinted · Shopify · Otto · Elsevier

Case Studies

Uber Eats - Scaling Search for Food Delivery · Airbnb - ML-Powered Experiences Ranking · Zalando - Self-DoS via Facet Aggregation · Slack - Enterprise Message Search with LTR · Etsy - Search Quality and Query Understanding · foodpanda - Classifying 300K Noisy Search Terms Across 16 Markets · Skyscanner - Learning to Rank for Flights · Netflix - Content Search Architecture · Canva - Search Pipeline Modernization · Vinted - Migrating Search from Elasticsearch to Vespa · Reddit - Vector Database Selection · Hybrid Fusion Failure - BM25 Displacing Reference Documents · Vespa - Ranking Without Labels on CORD-19

Videos

Conference talks and recorded presentations.

Max Irwin - The Search Engine Migration Circus — Haystack Live; search-engine migration playbook, “Hello Search”, feature parity, the “damage” metric & war stories

Rene Kriegler - Query Relaxation — OpenSource Connections; query relaxation as a recommendation problem, comparing heuristic, term-frequency, word2vec and neural approaches to predicting which query term to drop

Roman Grebennikov - Personalizing Search Results in Real-Time — Findify @ MICES 2019; real-time LTR personalization, position-bias feedback loops, shuffled exploration segments, purchase-weighted perfect rankings

Evgeniya Sukhodolskaya - Fine-Tuning Sparse Neural Retrievers for E-Commerce — Qdrant @ MICES 2026; why off-the-shelf SPLADE misfires on catalogs, the ANCE Hard Negative Mining loop, full vs inference-free SPLADE, specialize vs generalize

Evgeniya Sukhodolskaya - Relevance Feedback Inside the Search Engine — Qdrant @ Berlin Buzzwords 2026; index-native Relevance Feedback steering HNSW traversal, distilling a reranker into the index, and the case against black-box search engines

Key People

Daniel Tunkelang · Doug Turnbull · James Rubinstein · Omar Khattab · Jo Kristian Bergum · Trey Grainger · Andreas Wagner · Giovanni Fernandez-Kincade · Wolf Garbe · Eugene Yan · Andrew Kornilov

Stats

Counted 2026-08-05.

  • 314 article notes
  • 181 concept notes (incl. Ranking Signal Selection, Isolated Feedback Loops, Impression Bias, Out-of-Time Validation, Hashing Trick, Tokenization, Out-of-Vocabulary, Pooling, Score Normalization, Brute-Force Vector Search, Zero-Shot Retrieval, MUVERA, PCA, t-SNE, UMAP, TurboQuant, RaBitQ, BBQ, HNSW, SQ, BQ, Search-R1, Synthetic Query Generation, Consistency Filtering, Dense Passage Retriever, PROMPTAGATOR, FLAN-T5, Contrastive Learning, Staged Judging, Statistical Significance in Search Evaluation)
  • 59 topic notes (incl. Quepid Beyond Supported Engines, Multi-Tenancy in Search, Two-Sided Marketplace Ranking, Vector Search Tradeoffs, PCA vs t-SNE for Retrieval, Federated vs Unified Search, Migration between Search Engines, Elasticsearch Learning to Rank, Vespa Learning to Rank, Late Interaction in Vespa, Elasticsearch vs OpenSearch, Embedding Models Compared, Model Selection and Fine-Tuning Evaluation, Retrieval Benchmarks and Leaderboards)
  • 130 people notes (incl. Themis Mavridis, Soraya Hausl, Andrew Mende, Roberto Pagano, Jose Parreño, Roy Keyes, Davit Khachaturyan, Andrew Kornilov, Geoffrey Hinton, Laurens van der Maaten)
  • 14 case study notes (incl. Vespa - Ranking Without Labels on CORD-19, Hybrid Fusion Failure - BM25 Displacing Reference Documents, Vinted - Migrating Search from Elasticsearch to Vespa)
  • 47 company nodes (incl. Booking.com, Amazon Web Services, Elsevier)
  • 34 tool notes (incl. ann-benchmarks, Quepid, User Behavior Insights, Querqy, Elasticsearch, OpenSearch, Solr, Qdrant Vector DB, Weaviate Vector DB, FAISS)
  • 14 dataset notes (Amazon ESCI Dataset, BEIR, BRIGHT, ESCI-S Dataset, Home Depot Product Search Relevance, LoTTE, MIRACL, MS MARCO, MTEB, Natural Questions, RTEB, SIFT1M, TREC-COVID, WANDS Dataset)
  • 6 Maps of Content

See History for the full note-addition log.