A real incident
A shipped model quietly got worse, and the people who noticed had no way to prove it; the chapters here build the instrument that would have caught the regression before it reached anyone.
The chapters
58 min total
- The demo isn't the product7 minYou see the full range production sends: ordinary inputs, malformed ones, and the hostile tail.
- The quality bar: what good is7 minYou pin good to one page: the job, the load-bearing quality, and clearly scorable statements.
- Test cases: sample reality7 minYou build cases from real transcripts, tagged by input, user, stakes, split golden and living.
- Graders: score the outputs8 minYou match statements to the cheapest grader you can trust: code, a rubric judge, a human hour.
- The regression gate7 minYou make the set a ship rule: yesterday's score is the floor, and no change below it merges.
- Production signals7 minYou log signals users already send, retries and escalations, and turn each failure into a case.
- Goodhart's trap: gamed metrics7 minYou see why a measure goes bad once it becomes a target, and four defenses that keep it honest.
- Build your eval8 minYou assemble bar, cases, graders, and gate into one page, run weekly against what you ship.
The sum
Separately, the bar, cases, graders, and gate are just notes in different files. Run together on a schedule against what you ship, they become one instrument you use to catch a regression in private, before a customer catches it for you.
The Quality Bar is a one-page, fillable sheet. Fill it and you leave with your product's job, cases, graders, and floor score, ready for a weekly run.
The Quality BarFillable PDFDownload →