Benchmark
Cases decided before JUDAI saw them
The outcome was predicted from the facts alone, then checked against what the Court actually held. Predicting a decided case is not the product. It is the measure of whether the legal reasoning holds. Published whole: the wrong answers and the refusals included.
What this measures
A score with no architecture named is a claim, not a measurement. Each design is scored on its own, and the pair stays side by side.
Passages cut to 1,500 characters and chained through four summarising steps, so the model never saw a judgment whole. Replaced.
Whole documents read in full, then every quote matched against the source text before anything reaches the screen. This is the design answering questions today.
68% was measured on the same 24 cases, declining none.
That run was an experiment and was never stored, so it is not in the record below. This slot fills the day it is re-run and recorded, and then the two designs sit side by side on the same cases.
Where it was wrong
Ordered by how sure it was. The one at the top is the worst error a legal tool can make: confident and mistaken. It is here rather than behind a filter.
Calibration
When it says 80%, does it get 80% right? A system that knows what it does not know is usable even when it is often unsure.
Where it claimed 85% it was right 33% of the time, across 3 answers, too few to conclude much from, and published anyway.