Annabell Schäfer

Writes on the Langfuse blog about LLM evaluation and observability.


Contribution to This Vault

The argument worth carrying into relevance judging is about rubric design rather than cost. A judge that answers one atomic question at a time with a probability will return low confidence when the criteria are vague, instead of a crisp label that conceals the vagueness — so the tool surfaces a bad rubric rather than absorbing it. The piece is also clear-eyed about what the approach gives up: no rationale on any individual verdict, and no ability to abstain unless the question space includes an explicit escape hatch.

It is likewise careful with the benchmark it reports, noting that agreement with one frontier model is not accuracy, and that an open-source judge scored two points better for marginally more money.