Tag: evals
All the articles with the tag "evals".
The LLM judge is a model too
32 min readThe judge is a model too, with the same length, position and self-preference biases. Measuring it, versioning it, and knowing when not to use one.
Containment is measured on the calls that never needed you
36 min readContainment counts the sessions that never needed you. A booking agent's rollback is a refund, and a surge multiple is a scalar over a rotating mixture.
Faithful to the wrong version
26 min readOn a versioned corpus, faithfulness metrics pass an answer grounded in a revision retired months ago. Treating the corpus as bitemporal is what fixes it.