Benchmarks

One hundred rebuilt verdicts, scored against the judge.

100 US civil jury verdicts rebuilt from their public records, each scored by 250 simulated juries. A further 900 documented verdicts are listed below as a coverage catalog; those are not an accuracy measure and are excluded from the score.

direction, rebuilt suite
76/100
trial judge, same task
77.9%
when it commits
59/77
catalog, not scored
900
Real
P plaintiff, D defense, H hung.
Called
The side 250 simulated juries favored.
% plaintiff
Share of those juries, not a score.
Result
Match same side. Miss other side. Even is 45-55%.

1000/1000

CaseResult
How to read the score, including the misses

Direction on the rebuilt suite is 76 of 100. Even means the plaintiff share sat within five points of 50%, so the engine did not commit. On the stress rows where it does commit, the score is 59 of 77, and that is what the regression gate asserts on. Raw direction is not the gate: this suite is plaintiff-heavy, so a caller that always said plaintiff would score most of it without reading anything. We track per-class recall and calibration instead.

The 100 scored rows are full public-record rebuilds. The 900 catalog rows are matters in the shape the console creates - parties, claims, facts, damages and a trial plan from the public caption - but their feature vectors are inherited from one of 19 template cases chosen by program label, not derived from their own record. Most of a catalog row's call follows from that template, which is why those rows are listed for coverage and excluded from the score.

Real juries also split inside these programs. If even real juries split on near-identical records, the reliable move is a paired test: the same juries hear both versions of your case.

Compensatory damages on a single plaintiff are in scope. Punitive and multi-plaintiff awards are not. A row with no plaintiff share is marked not seated. It is never scored as a 0% defense miss.

Now run your case.

The same engine, on the record you are actually trying.