Line Lab Research

The Consensus: Three AIs Argue Over Every Bet We Fire. Can Their Score Predict Our P&L?

Our betting rules are deliberately dumb: mechanical triggers on ticket-vs-money divergence, no opinions allowed. The Consensus experiment asks whether adding an informed layer on top — actual pitching matchups, injuries, park factors, the stuff the splits can't see — carries any signal at all. We are not letting it touch money. We're making it show its work in public first.

Exactly what happens on every fire

When a Rule A/B/C trigger fires pre-game, a pipeline runs before first pitch (a leak guard refuses to score any game that has started — a score computed after results exist would be worthless):

The blind win probability matters because we can compare it to the price we'd actually pay on Polymarket. A panel that says 0.55 on a dog asking 0.46 is claiming nine points of edge; a panel that says 0.48 on a 0.46 ask is calling it a coin flip at fair price.

What would make this a success (and what wouldn't)

The primary test — fixed before the experiment started — is rank versus P&L, not win rate: after ~100 scored plays, do the high-scored fires make more money than the low-scored ones? If yes, the score becomes a sizing input. If no, we'll have bought the answer for $0, because none of this touches a bet. The gate lands around mid-September.

The live ledger

Every fire, its Polymarket ask at score time, the panel's blind probability, the score, and the outcome — updated at every nightly grade:

DateGameOur dog @ askPanel win probScoreResult
2026-08-04Twins@RoyalsRoyals @0.430.424/10WON
2026-08-04Athletics@RedsAthletics @0.460.444/10LOST
2026-08-02Royals@RockiesRoyals @0.460.558/10LOST
2026-08-02Cardinals@Blue JaysCardinals @0.450.517/10WON
2026-08-01Royals@RockiesRoyals @0.460.485/10LOST

Scored so far: 5 plays; graded record 2-3 (−0.45u at the recorded asks).

Since August 5 the same panel also scores two other populations — every remaining game on the board as a control arm, and the sides X handicappers take. Those are separate experiments on separate ledgers and they never touch the pre-registered pool above; they're written up in Crossing Two Experiments.

Honesty notes, same as everywhere on this site: the sample is tiny and means nothing yet — the whole point is the preregistered 100-play gate. The panel's scores are logged before results exist and never edited. No bet is placed, sized, or skipped because of this experiment while it's running. Nothing here is betting advice.

← All research