DEMO PREVIEW — Markets and probabilities shown are sample data, not real model output. Our ML engine is in development.

Get notified at launch →
Get notified at launch
← Back METHODOLOGY · CALIBRATION (PLANNED)

One real model, one planned engine. The Match Forecast Model v0.1 below is real — a calibrated football model with a genuine held-out reliability diagram. The broader event-contract engine described further down is still in development, and those category scores and figures are illustrative placeholders, clearly badged as such.

How we'll prove
we're honest.

When the engine is live, every probability will be graded against what actually happened. Categories where we beat the market will read blue; the ones we don't will sit honestly grey. This page previews that reporting — with example data for now.

MARKETS RESOLVED
0
OVERALL BRIER
STATUS
Pre-launch
MATCH FORECAST MODEL v0.1 real · held-out results

Our first model is real — and it's calibrated.

Built for the World Cup 2026, this model forecasts the 1X2 outcome of a single national-team match. The figures here are not placeholders: they come from a real, leakage-free walk-forward backtest on 8,000 held-out international matches. It is a calibrated probability estimate — not a betting signal.

RELIABILITY DIAGRAM · REAL
held-out 2018–2026
0% 0% 20% 20% 40% 40% 60% 60% 80% 80% 100% 100% MODEL PREDICTED PROBABILITY OBSERVED FREQUENCY
Every 1X2 probability the model produced on 8,000 held-out matches, binned: what it predicted (x) vs what actually happened (y). It hugs the diagonal — well calibrated — drifting slightly under-confident on heavy favourites (a known limit, below).
SCORE vs A NO-SKILL BASE RATE
MODEL BRIER 0.5135
BASE-RATE BRIER 0.6336
MODEL LOG-LOSS 0.8745
BASE-RATE LOG-LOSS 1.0507
vs MARKET not yet benchmarked

Lower is better. The model clearly beats a no-skill base rate (climatology). It has not yet been benchmarked against market odds — that pass is planned after the tournament.

How it works — Elo plus a draw model

National teams have no league, so the club models that price domestic football do not apply — you cannot compare a team's strength across competitions that never meet. The documented standard is a rating model, and we use World Football Elo: after every international, a team's rating moves toward the result, by more for a World Cup than a friendly, by more for a heavy win, with a home-advantage bump that switches off at neutral venues (where most World Cup games are played). The rating gap between two teams is their expected result.

Elo on its own has no notion of a draw, so we add the Davidson (1970) model — the standard way to turn a strength gap into three probabilities (home / draw / away) with a single draw-propensity parameter fit from the data. The result is the calibrated curve above: when we say 60%, it happens about 60% of the time.

What this model can't do (its limits, stated plainly)

  • Not yet benchmarked against the market. The free dataset we train on has no odds, so we have only shown the model beats a no-skill base rate — not that it matches or beats the market. Until that pass is done, treat the market column on the forecasts page as the reference, not our number.
  • Slightly under-confident on heavy favourites. Where it predicts ≥75%, the event happens a little more often than that — visible at the top-right of the diagram. A recalibration step can tighten it; v0.1 leaves it honest and unpatched.
  • Newcomers start average. A team with little international history begins at a neutral rating, so its earliest forecasts are the least reliable.
  • One draw parameter for all of football. The draw rate is a single global value; it doesn't yet flex by competition or era.

Full backtest, per-competition scores and the model card live in the engine repository. These are calibrated probability estimates for research and education — not financial advice, and not a betting signal.

LIVE TOURNAMENT TRACKING real · updates daily

Graded in public, as it happens.

Our World Cup 2026 page tracks the model live. To keep it honest we keep two separate forecast stores and never mix them:

PLAYED MATCHES → pre-tournament forecast

Forecasts are frozen before kickoff (ratings as of 2026-06-11) and never recomputed. Each result is scored with a multiclass Brier — so a played match becomes a live calibration datapoint that we cannot retro-fit.

UPCOMING MATCHES → current forecast vs market

Re-computed daily from every result so far, shown beside live Kalshi (CFTC-regulated) and Polymarket prices — reported separately, never averaged, with no "edge" or "beat the market" language.

LIVE TOURNAMENT BRIER · 24/72 PLAYED
0.6729 vs backtest 0.5135 · no-skill 0.6336

Running above the backtest right now — the opening round was upset-heavy and the sample is small. We publish it as-is; a model that only ever looks good is marketing, not research.

Liquidity quality, shown honestly

A price is only as good as the market behind it, so each venue's quote carries a liquidity tier. For Polymarket we use 24-hour volume (> $5k deep · $1k–5k moderate · < $1k thin); for Kalshi, which doesn't expose volume on these markets, we use the bid-ask spread (< 2c deep · 2–5c moderate · > 5c thin). On a thin market we say so rather than quote a price as if it were firm. The two venues are links, not affiliate placements — we earn no referral revenue from them.

THE PLANNED EVENT-CONTRACT ENGINE illustrative — in development
RELIABILITY DIAGRAM
Illustrative example
0% 0% 20% 20% 40% 40% 60% 60% 80% 80% 100% 100% MODEL PREDICTED PROBABILITY OBSERVED FREQUENCY
X-axis: what the model predicted. Y-axis: what actually happened. Perfect calibration sits on the dashed 45° line. Bubble size scales with sample count per bucket.
HOW IT WORKS, IN PLAIN ENGLISH
1 · Independent priors

For every contract, we price from scratch using ~140 features — never from the market price itself.

2 · Disagreement is signal

The gap between our prior and the market price is the edge. Sign tells direction, magnitude tells conviction.

3 · Honesty by construction

Every model is scored against realized outcomes. If we have no edge in a category, the UI shows it grey.

Our approach, in depth

MispriceHQ exists to answer one question for every event contract we cover: is the market price right? To answer it credibly, a model has to form its own opinion before it ever looks at the market. That is the first and most important design rule of the engine we're building.

1. Price from scratch, never from the market

For each contract, the planned model assembles features from primary data — rate curves and macro releases, options-implied paths, polling aggregates, on-chain crypto flows, and platform-level liquidity — and produces an independent probability. The current market price is deliberately excluded from the model's inputs. If we fed the market price in, the model would learn to echo the crowd, and the gap between the two would be meaningless. Keeping the prior independent is what makes disagreement informative.

2. Disagreement is the signal

We define edge as the difference between our model probability and the market price, in percentage points. The sign tells you direction — whether we think the event is more or less likely than the market does — and the magnitude is a rough proxy for conviction. A contract trading at 34% that our model prices at 51% carries a +17pp edge. Edge is not a promise; it is a hypothesis that gets graded the moment the market resolves.

3. Honesty by construction

Most forecasting products show you only their hits. We intend to do the opposite: publish the score for every category, including the ones where we have no edge. When a category's model cannot beat the market baseline out-of-sample, the interface will say so — the markets there will sit visibly grey rather than glowing with false confidence. The boring, honest grey is the point. A research tool that only ever looks confident is not a research tool; it is marketing.

Why we lead with calibration, not win rate

A forecaster who says “70%” and is right 70% of the time is doing excellent work, even though they are “wrong” 30% of the time. Win rate — how often the higher-priced side happens — rewards confident guessing and punishes honest uncertainty. It is the wrong yardstick for probabilistic forecasts. The right yardstick is calibration: when we say 70%, does the thing happen about 70% of the time?

The standard way to measure this is the Brier score — the mean squared error between the predicted probability and the realized outcome (1 for yes, 0 for no). Lower is better; 0 is perfect, 0.25 is what you'd get by guessing 50% on everything. We grade against the market-consensus baseline in the same category, so the only number that matters is whether our Brier score beats the price you could have just taken off the screen. We explain this in depth in our guide to why Brier score beats win rate.

Calibration is also why the reliability diagram above is the first chart on this page. When the engine is live, every published probability becomes a dot on that chart the day its market resolves. Points on the dashed 45-degree line are perfectly calibrated; points above mean we were under-confident, points below mean we were over-confident. There is nowhere to hide on a reliability diagram, which is exactly why we use it.

What you should expect at launch

We would rather ship late and honest than early and wrong, so a few commitments are worth stating now:

  • Sample data until it's real. Everything you see on the site today is placeholder data for design purposes, marked as such on every surface. We will not quietly swap in “live” numbers — the launch of real model output will be explicit.
  • Publish only where we have edge. A category goes live only after its model beats the market baseline on held-out, already-resolved markets. Categories that don't clear that bar stay in research and are labeled accordingly.
  • Every number ships with its calibration. Probabilities will always be accompanied by the category's track record, sample size, and confidence interval — never a bare figure.
  • We disclose our limits. Base rates in some categories are genuinely noisy; regimes shift; tails happen. Where our confidence interval is wide, we'll show it wide.

Until then, the most useful thing on this site is the published methodology you're reading and the research library. Both are written to the same standard the models will be held to: sourced, dated, and honest about uncertainty.

Nothing on this page is financial advice. Event contracts are derivatives that carry risk of total loss. Model accuracy, once published, will describe past performance and is not a guarantee of future results.

SCORE BY CATEGORY
Lower Brier is better. At launch, categories where our score beats the market consensus will be highlighted. Values below are an illustrative example, not real results.
Illustrative example
Category Model Brier Market Brier Δ Edge N Verdict
Macro 0.182 0.211 -0.029
312 ● EDGE
Crypto 0.196 0.224 -0.028
248 ● EDGE
Tech 0.218 0.222 -0.004
174 ○ NO EDGE
Politics 0.209 0.215 -0.006
142 ○ NO EDGE
Culture 0.226 0.234 -0.008
100 ○ NO EDGE
Sports 0.241 0.232 +0.009
96 ○ NO EDGE — WORSE
Sports is currently worse than the market. We leave it visible on purpose — when we don't have edge, the UI says so. We are not running models on sports markets in production.
PLANNED DATA SOURCES
FRED — macro time series rate curves, CPI, unemployment, PCE
BLS / BEA public releases labor & GDP first-prints
CME options chain Fed funds futures, implied path
Spot & on-chain crypto feeds ETF flows, exchange balances
Public polling aggregators topline + crosstabs, weighted
Platform order books Polymarket, Kalshi, Manifold
MODEL CARD · PLANNED
ARCHITECTURE GBM ensemble + isotonic calibration
TARGET FEATURES ~140
PLANNED REFIT NIGHTLY
HOLDOUT 30-DAY ROLLING
NEVER USED AS INPUT MARKET PRICE
NOT FINANCIAL
ADVICE.

MispriceHQ publishes independent probability estimates for educational and research purposes. Models can be wrong; market microstructure can shift overnight; tails do tail. We don't take positions, we don't route orders, and we are not your fiduciary. Risk of total loss is real and material.

Event contracts on Kalshi are regulated derivatives traded under CFTC oversight; other venues operate under their own jurisdictions. By using this site you acknowledge our methodology, our calibration, and that the numbers above describe the past — not the future.