TL;DR: it breaks an LLM judge run into claims->evidence->verdicts and flags when a verdict is not supported by the evidence, so i can check it manually
Show HN: I made a small helper for checking model-graded answers
- Posted 2 hours ago by ML0037
- 1 points