History — 2026 week 34 (Aug 17 – Aug 23, 2026)
Newest first.
2026-08-21 — Tuning BM25 for E-commerce Search (1 new, 10 updated)
BM25’s k1/b defaults are tuned for prose, not product catalogs — six-word titles, single-token brands, and merchant descriptions padded with keyword variants behave differently from ordinary documents. Tuning BM25 for E-commerce Search argues for field-by-field parameters, low k1/b on title, brand and category, defaults left alone on description, and grounds the payoff in Shopify’s retuning of its product-title field via Bayesian Optimization. It adds a second reason specific to marketplaces: sellers who control their own listings will discover that repeating query terms or shortening titles improves ranking, so a lower k1 on seller-controlled fields blunts a manipulation incentive, not just a relevance tweak.
Topics — Tuning BM25 for E-commerce Search Updated — BM25 · Bayesian Optimization · E-commerce Search · Two-Sided Marketplace Ranking · Haystack US 2022 - Bayesian Optimization of Relevance at Shopify · Doug Turnbull · Andy Toulis · Max Irwin · Shopify · Quepid
2026-08-21 — Bayesian Optimization of BM25 at Shopify (3 new, 5 updated)
Bayesian Optimization now covers the technique Doug Turnbull frames as a lightweight halfway point between hand-tuned boosts and full Learning to Rank: a surrogate model over past (parameters, NDCG) observations replaces grid search when retuning BM25 k1/b and other ranking-function parameters. Haystack US 2022 - Bayesian Optimization of Relevance at Shopify walks through Andy Toulis’s worked example at Shopify: a first optimization run exploited presentation bias baked into the training data before a constrained rerun flattened BM25’s length-normalization curve on product titles and validated the gain on held-out data. Talk given at Haystack US 2022, hosted by OpenSource Connections.
Videos — Haystack US 2022 - Bayesian Optimization of Relevance at Shopify Concepts — Bayesian Optimization People — Andy Toulis Updated — Doug Turnbull · Shopify · OpenSource Connections · Haystack US · BM25
2026-08-20 — Allegro’s LLM-Judge Relevance Framework (4 new, 6 updated)
Few-shot prompting is supposed to help an LLM as Judge — Automating Search Relevance Assessment at Scale with LLM-as-a-Judge found the opposite: stripping examples from the prompt raised both accuracy and inter-rater agreement, while structured step-by-step reasoning instructions did the real work. The note covers Allegro’s Relevance Assessment Tool, validated against a 380K-pair dataset spanning Polish, Czech, Slovak and Hungarian catalog search, and its migration of batch judging from a cloud model to a locally-hosted 26B Gemma variant — cutting inference cost 60% at parity with the cloud baseline on Polish, while “no-thinking” inference beat chain-of-thought reasoning outright. Sourced from the Allegro Tech engineering blog, authored by Joanna Marhula and Mateusz Sidor.
Articles — Automating Search Relevance Assessment at Scale with LLM-as-a-Judge People — Joanna Marhula · Mateusz Sidor Companies — Allegro Updated — LLM as Judge · Search Quality Assurance · Multilingual Search · Judgment Lists · NDCG · Model Selection and Fine-Tuning Evaluation