Rubricon

Scoring harness for LLM-generated conversation feedback. A criterion score only stands if the grader can cite the part of the transcript it is talking about, and the citation is checked against the transcript rather than trusted.

rubric Hard conversations, v1 version 6b5ad108c775fb03 grader gemini-2.5-flash

Transcript

Try the third button. It runs the same grader but swaps one criterion's citation for a fluent quote that never appears in the transcript. That is what a model does when it justifies a score it cannot support. The evaluation is refused rather than reported, which is the assertion in refuses a score whose citation is not in the transcript.