Calibration
Every call we make, scored against the market price on the same question. The Brier score measures squared error — lower is better — and it punishes confident misses hardest. This is the public measurement of whether the machine beats the crowd.
28 verdict-bearing calls with both stamps, of 55 resolved. Row-by-row: EHIQ closer on 14, market closer on 14, ties 0.
By venue
| venue | n | EHIQ | market | delta |
|---|---|---|---|---|
| polymarket | 21 | 0.2130 | 0.2212 | -0.0082 |
| kalshi | 5 | 0.3031 | 0.2588 | +0.0443 |
| manual | 1 | 0.3025 | 0.1225 | +0.1800 |
| sell_side_consensus | 1 | 0.3136 | 0.1225 | +0.1911 |
day-0 baseline 2026-09-13: the machine-proposal batch (#371-376) resolves inside the next 60 days; this curve is where "is the machine beating the crowd" gets answered.
How the outcome is recovered: hit=True means the call was right, not that the event happened — a 2% call that “hits” is the event not happening. Outcome = (stamp ≥ 0.5) == hit, exact at the 0.5 threshold. The first readout of this page scored 0.42 by forgetting that; the rule is printed here so it never happens again.