Research · October 11, 2026

Nothing we built beats “the petitioner wins”

Every number below is BACKTESTED on known outcomes.

At the Supreme Court, the side that asked the Court to take the case usually wins. Three instruments built to read oral argument were tested against that one-line rule. Two lost to it and the third tied it by copying it.

SCOTUSblog’s Adam Feldman reported the same pattern for OT2025 in July: across 56 decisions, no oral-argument rule beat always picking the petitioner. This scorecard runs the two terms before it, OT2023 and OT2024, under a ladder written down before the first run, with paired tests and a shuffled-label control.

1. What we tested, and what we froze first

The test was written down on September 28, before any rung on it was run. It set a ladder of five rungs and a rule for losing.

  • Rung 1, the floor: always pick the petitioner.
  • Rung 2: the side that draws more questions from the bench tends to lose (the Epstein, Landes and Posner question rule).
  • Rung 3: the same rule, counted in the words of the justices’ questions instead of the number of questions.
  • Rung 4: a separate model for each justice, reading how that justice’s questions and words split between the two sides, then a majority of the nine.
  • Rung 5, the ceiling: 74.04%, the benchmark our pre-registration set as the top of the plausible range. Anything above it starts a search for data leaking into the test before anyone celebrates.

The losing rule was fixed in advance: if the rungs cannot beat “always petitioner”, they do not ship, and the result prints at full size. Rung 4 got its own written spec on September 29, frozen before its first fit, with a shuffled-label control that had to fail.

The instrument reads oral argument only. Briefs, opinions and lower-court records are separate material and are never mixed in. Each term is graded on its own and never pooled, and no comparison is claimed on fewer than eight cases.

2. The scorecard

Rungs 1 to 3, every graded case per term:

TermCasesRung 1: always petitionerRung 2: question ruleRung 3: word rule
OT20235877.6%46.6% (p=0.001)22.4% (p=0.000)
OT20244879.2%46.8% (p=0.006, 47 cases)20.8% (p=0.000)
OT20252100.0%50.0%50.0%

The p-values are McNemar tests against rung 1 on the same cases. OT2025 has 2 graded cases, so it supports no claim.

Rung 4, walk-forward (trained on earlier terms, graded on the next):

RunCasesRung 4Rung 1 on the same casesMcNemar
Train OT2023, grade OT20244168.3%68.3%p=1.000
Train OT2023+24, grade OT2025475.0%75.0%p=1.000 (no claim, n<8)
Control: shuffled labels, OT20244168.3%68.3%p=1.000

Rung 4 picked the petitioner in all 41 graded cases. With the labels scrambled it scored exactly the same. Its 41 cases are the subset with joined vote records, which is why its floor reads 68.3% and not the 79.2% of the full term. Rung 4 also takes its outcomes from the vote records, which is why it grades 4 OT2025 cases where the rungs above grade 2.

All three rungs above the floor land at or below it. The question and word rules lose to it decisively in both graded terms. The per-justice model ties it by never disagreeing with it once.

3. What the counts show

Counted in questions, the split is near even: the more-questioned side won 31 of 58 OT2023 cases and 25 of 47 OT2024 cases. Counted in words, the bench put more on the petitioner’s side in 56 of 58 OT2023 cases and 46 of 48 OT2024 cases, whoever won. So the word rule votes against the petitioner almost every time, and the petitioner usually wins.

One counting rule bears on that word result. Both counts credit the petitioner with the justices’ turns during rebuttal, because rebuttal is the petitioner’s time at the podium. The word imbalance may partly reflect that rule. Neither count separates winners from losers better than knowing who filed the petition.

The per-justice model is pulled toward each justice’s own record, and in this corpus that record leans to the petitioner. The question and word imbalances never move a justice far enough to vote the other way, so every justice picks the petitioner and the majority follows. The shuffled-label control shows what is left: the base rate.

The same result has come back before. In August two earlier models lost to the base rate by 11.1 and 21.4 points, and the holdout recorded minus 31.2. Every oral-argument instrument built or imported here either rides the petitioner rate or loses to it.

4. The 78% is this corpus’s number

Published rates run lower. The Court has reversed about 71% of the cases it decided since 2007, by Ballotpedia’s count, and petitioners won 67.9% of OT2025’s signed decisions in SCOTUSblog’s count. This corpus sits near 78–79% because it is a selected set: the transcripts on hand, the outcomes that could be coded, and vacated judgments counted as petitioner wins. That makes the corpus’s own rate the only fair benchmark for anything graded on it, and every comparison here uses it.

Rung 1 sits above the 74.04% ceiling, and that alarm points at the corpus: rung 1 is the corpus’s own base rate.

5. What the instrument may say about Suncor and Intel

Suncor Energy v. County Commissioners of Boulder County was argued on October 5 and Anderson v. Intel Corp. Investment Policy Committee on October 6. On cases like these, the instrument’s permitted output is narrow:

  • A tone read, published as a band across repeated scorer passes, never a single point.
  • Family-level description only, such as “questioning ran against the petitioner”. No pattern or type names until our labels pass an agreement test.
  • Question counts are the baseline a tone read has to beat.
  • No outcome forecast. Any call on either case goes through the dated public ledger with its market declared at creation. If no venue lists the case, it is graded on its own, with no market price attached later.

6. What comes next

Acoustic features (pitch, interruptions, laughter) are a different class of signal and remain untested here. They already have their own written spec, dated September 29 and frozen before any audio is touched, and they start from this result. The bar is the same: beat “the petitioner wins” on a term the model has not seen, or print the loss at full size.


Sources

← All research

Backtested research on public court records. Not legal or investment advice.