Gloss — Evaluation set

Frozen review routes for evaluating generated companion notes. The existing 25-note review and seven-note checking set are retained in the linked reports.

Score each factor as supported, related-but-misleading or unsupported, and record missed broad facets separately. Review dates, polarity, procedural constraints and specific names explicitly.

Reports: selection experiment, evidence gate, and experiment ranking.

This is an exploratory set, not blind gold-standard data. Add every production false match as a new labelled case before changing the reranker.

version 1  ·  created 2026-09-18  ·  updated 2026-09-18  ·  tags gloss, evaluation, draft