Evidence
The record, including the parts that are not flattering.
Every edge the engine flags is logged as a forward-tested bet before anyone is asked to pay for it, then graded against the price the same book closed at. This is that record. The error bars are here because a return without one is not evidence, and at this sample they are wider than anything inside them.
Against the closing line
Beating the close is the only measure that predicts long-run profit rather than describing a streak. Only same-book closings in markets we have verified count here: 0 bets sit in markets still proving themselves and 0 could only be graded against a different book, and neither is allowed to move this number.
Edges vs the close
3.96%
95% CI 3.63% to 4.31%
Beat the close
69.18%
above the 55% we hold ourselves to
Graded bets
1577
past the 150 we treat as a floor
Does a bigger edge beat the close by more?
If the model is measuring something real, the edges it rates highest should beat the close by the most. If the bars are flat or backwards, the ratings are decoration.
Flagged 2-3%
3.69%
n=1005
Flagged 3-4%
4.03%
n=339
Flagged 4-5%
4.71%
n=167
Flagged above ceiling
5.80%
n=66
Higher flagged EV has beaten the close by more, in order, across every band with data.
Why the board stops at 5%
Most tools like this will show you a 15% edge. We used to. On 29 July we measured every bet we had ever flagged against the probability our own model gave it, and the answer was that our estimate is honest up to about 4% and stops being honest above it. So the board stops calling anything above 5% an edge.
Flagged 2-3%
z +0.14
n=313, the model is right here
Flagged 5%+
z −3.43
n=52, ROI −48.9%
Volume this costs
6.9%
of everything we flagged
The number is a standard-score: how far the results landed from what we predicted, in standard deviations. Near zero means the model told the truth. At 5% and above it is more than three deviations wrong, and in the same direction every time.
Nothing is hidden. Prices above the ceiling stay on the board with the reason written on them, and we keep grading every one of them, so if this decision is wrong the record will say so rather than the ceiling quietly protecting itself.
What it would have done to a bankroll
The question a buyer actually asks, and a harder one to answer honestly, so it is never used as a gate.
Flat-stake return
-0.47%
plus or minus 2.28% on 2175 settled
Quarter Kelly
-43.44%
worst drawdown 57.53% peak to trough
Distance from zero
0.20σ
inside the noise band
At this sample the return is consistent with no edge and with a good one. Both are true, which is why the figure above is not printed in a colour. A real 3% edge would need about 5,041 settled bets to clear zero by two standard errors. We have 2175. Closing-line value converges far faster than realised return, which is why the number above it is the one we are judged on.
Market by market
A market is marked verified only after its own record passes the same test. Until then it is shown in the product with a badge and kept out of the headline above.
2350 forward-tested bets in total, 1577 with a closing line so far, plus 6463 candidates the engine declined and graded anyway so the decision can be judged later.
Read it yourself
The sample market runs the same engine on a recorded board, with no account and no card. Every number on it is produced the way the numbers above were.
Open the sample market