One real model, one planned engine. The
Match Forecast Model v0.1 below is real —
a calibrated football model with a genuine held-out reliability diagram. The broader
event-contract engine described further down is still in development, and those category scores
and figures are illustrative placeholders, clearly badged as such.
How we'll prove we're honest.
When the engine is live, every probability will be graded against what actually happened.
Categories where we beat the market will read blue; the ones we
don't will sit honestly grey. This page previews that reporting — with example data for now.
MARKETS RESOLVED
0
OVERALL BRIER
—
STATUS
Pre-launch
MATCH FORECAST MODEL v0.1real · held-out results
Our first model is real — and it's calibrated.
Built for the World Cup 2026, this model forecasts the 1X2
outcome of a single national-team match. The figures here are not placeholders:
they come from a real, leakage-free walk-forward backtest on 8,000 held-out
international matches. It is a calibrated probability estimate — not a betting signal.
RELIABILITY DIAGRAM · REAL
held-out 2018–2026
Every 1X2 probability the model produced on 8,000 held-out matches,
binned: what it predicted (x) vs what actually happened (y). It hugs the diagonal — well
calibrated — drifting slightly under-confident on heavy favourites (a known limit, below).
SCORE vs A NO-SKILL BASE RATE
MODEL BRIER0.5135
BASE-RATE BRIER0.6336
MODEL LOG-LOSS0.8745
BASE-RATE LOG-LOSS1.0507
vs MARKETnot yet benchmarked
Lower is better. The model clearly beats a no-skill base rate (climatology). It has not yet
been benchmarked against market odds — that pass is planned after the tournament.
How it works — Elo plus a draw model
National teams have no league, so the club models that price domestic football do not apply —
you cannot compare a team's strength across competitions that never meet. The documented
standard is a rating model, and we use World Football Elo:
after every international, a team's rating moves toward the result, by more for a World Cup
than a friendly, by more for a heavy win, with a home-advantage bump that switches off at
neutral venues (where most World Cup games are played). The rating gap between two teams is
their expected result.
Elo on its own has no notion of a draw, so we add the Davidson (1970) model —
the standard way to turn a strength gap into three probabilities (home / draw / away) with a
single draw-propensity parameter fit from the data. The result is the calibrated curve above:
when we say 60%, it happens about 60% of the time.
What this model can't do (its limits, stated plainly)
Not yet benchmarked against the market. The free dataset we train on has no
odds, so we have only shown the model beats a no-skill base rate — not that it
matches or beats the market. Until that pass is done, treat the market column on the
forecasts page as the reference, not our number.
Slightly under-confident on heavy favourites. Where it predicts ≥75%, the
event happens a little more often than that — visible at the top-right of the diagram. A
recalibration step can tighten it; v0.1 leaves it honest and unpatched.
Newcomers start average. A team with little international history begins at
a neutral rating, so its earliest forecasts are the least reliable.
One draw parameter for all of football. The draw rate is a single global
value; it doesn't yet flex by competition or era.
Full backtest, per-competition scores and the model card live in the engine repository. These
are calibrated probability estimates for research and education — not financial advice, and not
a betting signal.
LIVE TOURNAMENT TRACKINGreal · updates daily
Graded in public, as it happens.
Our World Cup 2026 page tracks the model live. To keep it honest we
keep two separate forecast stores and never mix them:
PLAYED MATCHES → pre-tournament forecast
Forecasts are frozen before kickoff (ratings as of 2026-06-11) and
never recomputed. Each result is scored with a multiclass Brier — so a played match becomes a live
calibration datapoint that we cannot retro-fit.
UPCOMING MATCHES → current forecast vs market
Re-computed daily from every result so far, shown beside live Kalshi
(CFTC-regulated) and Polymarket prices — reported separately, never averaged,
with no "edge" or "beat the market" language.
LIVE TOURNAMENT BRIER · 24/72 PLAYED
0.6729vs backtest 0.5135 · no-skill 0.6336
Running above the backtest right now — the opening round was upset-heavy and the sample is small. We publish it as-is; a model that only ever looks good is marketing, not research.
Liquidity quality, shown honestly
A price is only as good as the market behind it, so each venue's quote carries a liquidity tier. For
Polymarket we use 24-hour volume (> $5k deep · $1k–5k moderate · < $1k thin);
for Kalshi, which doesn't expose volume on these markets, we use the bid-ask spread
(< 2c deep · 2–5c moderate · > 5c thin). On a thin market we say so rather than quote a price as
if it were firm. The two venues are links, not affiliate placements — we earn no
referral revenue from them.
THE PLANNED EVENT-CONTRACT ENGINEillustrative — in development
RELIABILITY DIAGRAM
Illustrative example
X-axis: what the model predicted. Y-axis: what actually happened. Perfect calibration sits on
the dashed 45° line. Bubble size scales with sample count per bucket.
HOW IT WORKS, IN PLAIN ENGLISH
1 · Independent priors
For every contract, we price from scratch using ~140 features — never from the market
price itself.
2 · Disagreement is signal
The gap between our prior and the market price is the edge. Sign tells direction,
magnitude tells conviction.
3 · Honesty by construction
Every model is scored against realized outcomes. If we have no edge in a category, the
UI shows it grey.
Our approach, in depth
MispriceHQ exists to answer one question for every event contract we cover: is the
market price right? To answer it credibly, a model has to form its own opinion before
it ever looks at the market. That is the first and most important design rule of the engine
we're building.
1. Price from scratch, never from the market
For each contract, the planned model assembles features from primary data — rate curves and
macro releases, options-implied paths, polling aggregates, on-chain crypto flows, and
platform-level liquidity — and produces an independent probability. The current market price is
deliberately excluded from the model's inputs. If we fed the market price in, the model would
learn to echo the crowd, and the gap between the two would be meaningless. Keeping the prior
independent is what makes disagreement informative.
2. Disagreement is the signal
We define edge as the difference between our model probability and the market
price, in percentage points. The sign tells you direction — whether we think the event is more
or less likely than the market does — and the magnitude is a rough proxy for conviction. A
contract trading at 34% that our model prices at 51% carries a +17pp edge. Edge is not a
promise; it is a hypothesis that gets graded the moment the market resolves.
3. Honesty by construction
Most forecasting products show you only their hits. We intend to do the opposite: publish the
score for every category, including the ones where we have no edge. When a category's model
cannot beat the market baseline out-of-sample, the interface will say so — the markets there
will sit visibly grey rather than glowing with false confidence. The boring, honest grey is the
point. A research tool that only ever looks confident is not a research tool; it is marketing.
Why we lead with calibration, not win rate
A forecaster who says “70%” and is right 70% of the time is doing excellent work,
even though they are “wrong” 30% of the time. Win rate — how often the higher-priced
side happens — rewards confident guessing and punishes honest uncertainty. It is the wrong
yardstick for probabilistic forecasts. The right yardstick is calibration:
when we say 70%, does the thing happen about 70% of the time?
The standard way to measure this is the Brier score — the mean squared error
between the predicted probability and the realized outcome (1 for yes, 0 for no). Lower is
better; 0 is perfect, 0.25 is what you'd get by guessing 50% on everything. We grade against the
market-consensus baseline in the same category, so the only number that matters is whether our
Brier score beats the price you could have just taken off the screen. We explain this in depth
in our guide to
why Brier score beats win rate.
Calibration is also why the reliability diagram above is the first chart on this page. When the
engine is live, every published probability becomes a dot on that chart the day its market
resolves. Points on the dashed 45-degree line are perfectly calibrated; points above mean we
were under-confident, points below mean we were over-confident. There is nowhere to hide on a
reliability diagram, which is exactly why we use it.
What you should expect at launch
We would rather ship late and honest than early and wrong, so a few commitments are worth
stating now:
Sample data until it's real. Everything you see on the site today is
placeholder data for design purposes, marked as such on every surface. We will not quietly
swap in “live” numbers — the launch of real model output will be explicit.
Publish only where we have edge. A category goes live only after its model
beats the market baseline on held-out, already-resolved markets. Categories that don't clear
that bar stay in research and are labeled accordingly.
Every number ships with its calibration. Probabilities will always be
accompanied by the category's track record, sample size, and confidence interval — never a
bare figure.
We disclose our limits. Base rates in some categories are genuinely noisy;
regimes shift; tails happen. Where our confidence interval is wide, we'll show it wide.
Until then, the most useful thing on this site is the published methodology you're reading and
the research library. Both are written to the same standard the models
will be held to: sourced, dated, and honest about uncertainty.
Nothing on this page is financial advice. Event contracts are derivatives that carry risk of
total loss. Model accuracy, once published, will describe past performance and is not a
guarantee of future results.
SCORE BY CATEGORY
Lower Brier is better. At launch, categories where our score beats the market consensus
will be highlighted. Values below are an illustrative example, not real results.
Illustrative example
Category
Model Brier
Market Brier
Δ
Edge
N
Verdict
Macro
0.182
0.211
-0.029
312
● EDGE
Crypto
0.196
0.224
-0.028
248
● EDGE
Tech
0.218
0.222
-0.004
174
○ NO EDGE
Politics
0.209
0.215
-0.006
142
○ NO EDGE
Culture
0.226
0.234
-0.008
100
○ NO EDGE
Sports
0.241
0.232
+0.009
96
○ NO EDGE — WORSE
Sports is currently worse than the market. We leave it
visible on purpose — when we don't have edge, the UI says so. We are not running models on
sports markets in production.
PLANNED DATA SOURCES
FRED — macro time series rate curves, CPI, unemployment, PCE
BLS / BEA public releases labor & GDP first-prints
Public polling aggregators topline + crosstabs, weighted
Platform order books Polymarket, Kalshi, Manifold
MODEL CARD · PLANNED
ARCHITECTUREGBM ensemble + isotonic calibration
TARGET FEATURES~140
PLANNED REFITNIGHTLY
HOLDOUT30-DAY ROLLING
NEVER USED AS INPUTMARKET PRICE
NOT FINANCIAL ADVICE.
MispriceHQ publishes independent probability estimates for educational and research purposes.
Models can be wrong; market microstructure can shift overnight; tails do tail. We don't take
positions, we don't route orders, and we are not your fiduciary.
Risk of total loss is real and material.
Event contracts on Kalshi are regulated derivatives traded under CFTC oversight; other
venues operate under their own jurisdictions. By using this site you acknowledge our
methodology, our calibration, and that the numbers above describe the past — not the future.