01
A score anyone can re-derive
Chose A Python score computed from per-criterion verdicts over letting the model write its own score out of 100.
A score scraped from the model’s prose couldn’t be audited or explained. Now it’s arithmetic on verdicts and evidence that anyone can re-check.
Trade-off: Rubrics, verdict bands and weights became code to maintain, covered by 113 tests.
02
Evidence, or it counts for less
Chose Fuzzy matching of every quote a verdict cites over trusting the evidence the model says it found.
Strong verdicts often arrived with no evidence at all. Now a verdict whose quote isn’t in the submission or research is discounted.
Trade-off: An honest paraphrase fails the match too, so it’s discounted along with invented quotes.
03
Same judges for every revision
Chose The original reviewer panel for every revision over drawing a fresh panel on each run.
Each reviewer re-weights the rubric, so a new panel changes the maths and the score change stops meaning better or worse.
Trade-off: A pivoted idea keeps reviewers that no longer fit, so users got an opt-out to draw a new panel.