Accuracy scorecard

Every contract W.E.T. carries is a binary question that eventually resolves YES or NO at its venue, so the prices can be scored against what actually happened. This page reports exactly one thing: the calibration of the 428 contract prices the daily slate-basket mark tape has so far seen priced while open and later observed resolving. Which contracts those are, and which they are not, is stated before any number.

What this sample is — and is not

It is the slate-basket mark tape. The daily slate-basket mark tape — the two-sided price recorder that marks every holding in every published W.E.T. slate basket, once per session. Across 48 sessions (2026-08-01 → 2026-09-23) it covers 110 of the 135 slate books on record — counted by checking each book’s holdings against the tape, not by counting books — and 428 of those holdings have resolved.

It is not the governed benchmarks — WETGRI, WETFED, WETX and WETFRAG. Benchmark constituents that have resolved are scored separately, below.

And no index level is ever scored. Nothing that happens later makes a printed level right or wrong — an index reads prices, it does not forecast anything. That holds permanently, for benchmarks and slates alike, however large this sample grows.

By category
  • sports168
  • climate106
  • companies37
  • elections26
  • commodities20
  • crypto16
  • economics16
  • culture16
  • politics11
  • tech-science11
  • financials1

Read off the category frozen on each holding when the basket opened it, not inferred here.

By slate basket
  • storm-watch106
  • soccer-heat104
  • favorites-combo78
  • sports-heat71
  • earnings-beat37
  • gridiron-futures25
  • election-week21
  • commodity-stress20
  • awards-race16
  • crypto-greed12
  • ai-race11
  • inflation-watch11
  • flashpoint-watch9
  • election-heat7
  • geopolitical-risk5
  • bitcoin-path4
  • climate-risk4
  • culture-heat3
  • equity-greed3
  • global-easing3
  • midterm-control3
  • washington-week3
  • altcoin-heat2
  • recession-odds2
  • seats-elections-inside-two-months2
  • anthropic-financials-inside-six-months1
  • capture-politics-inside-six-months1
  • fed-path1
  • hike-economics-inside-six-months1
  • hike-economics-inside-six-months-inverse1
  • meeting-economics-inside-six-months1
  • meeting-economics-inside-two-months1
  • openai-financials-inside-six-months1
  • rate-economics-inside-six-months1
  • rate-economics-inside-six-months-inverse1
  • rate-economics-inside-two-months1
  • released-tech-science-inside-six-months1
  • seats-elections-inside-two-weeks1

A contract can sit in several baskets, so these are memberships rather than a partition and can sum above 428. All 428 observations are accounted for here.

  • ·8 benchmark constituents have also been observed resolving; they are scored separately in the benchmark cohort and are excluded from the counts above so no observation is counted twice.
The sample · Slate basket contract prices
428graded outcomes· resolution observed 2026-08-022026-09-23
Resolved YES 265 · NO 163
Venues polymarket 317 · kalshi 99 · gemini 12
Horizon at most 10d · tape 48d
Log-loss clamp bound on 0 rows

Every contract held in a published W.E.T. SLATE basket that the daily mark tape saw priced while the question was open and later observed resolving at its venue. Slate baskets only — it excludes the governed benchmarks (WETGRI, WETFED, WETX, WETFRAG), whose constituents are covered by the separate benchmark cohort.

Every figure below describes those 428 contracts and nothing else — not the benchmarks, not the desk as a whole, and not any position or result.

Brier score
0.1428
n = 428 · min 30

Mean squared error of the price against the outcome. Lower is better. A constant 50% forecast scores 0.25.

Brier skill vs base rate
+0.394
n = 428 · min 30

Against a forecaster who knew only how often these contracts resolve YES. Positive means the prices carried real information; a raw Brier alone cannot tell you this.

Base rate
61.9%
n = 428 · min 30

Share of the sample that resolved YES. The reference forecast the skill score is measured against.

Log loss
0.4231
n = 428 · min 30

Secondary metric, clamped to stay finite. It punishes confident misses far harder than the Brier, which makes it noisy on small samples.

Calibration by decile

The table a single score cannot fake: a systematic bias hides inside an average but not inside ten rows. A bucket publishes its observed frequency only once it holds 30 graded outcomes. Below that the value is withheld and the row says how many more it needs — an observed frequency built from six contracts is ±20 points of binomial noise, which is larger than any miscalibration worth reporting.

Priced atContractsMean pricedResolved YESObserved freq.Gap
0–10%
461.7%00.0%-1.7pp
10–20%
914.5%2needs 21 more
20–30%
524.8%0needs 25 more
30–40%
2936.3%10needs 1 more
40–50%
4045.1%2152.5%+7.3pp
50–60%
7754.2%4355.8%+1.6pp
60–70%
5263.6%3261.5%-2.1pp
70–80%
3674.7%3083.3%+8.7pp
80–90%
4184.7%3687.8%+3.1pp
90–100%
9396.5%9197.9%+1.4pp
7/10 buckets publishing · expected calibration error withheldfull table ≈ 300 outcomes
What is wrong with this sample
  • ·Horizons run from at most 1 to at most 10 days. Calibration is not constant across horizon, so these are pooled, not comparable across the range.

What was dropped, and why

The scored count is a subset of everything the tape saw close. Publishing the rejects is the only way to show the sample was not selected after the fact — so rows 2–5 partition row 1 exactly, and the last row splits the scored contracts between the two cohorts on this page.

Contracts the tape saw close490The denominator. Distinct contracts that reached any terminal state — resolved YES, resolved NO, or voided.
Scored436A live price exists on an earlier tape day than the day the answer was observed.
Dropped — voided15Closed by the venue with no gradeable answer. Scoring a void as NO would punish a price for a question never asked.
Dropped — already resolved when first seen0We never observed a live price. Pairing the settlement mark with itself would drive the Brier to zero — the classic way a fake scorecard looks superb.
Dropped — no live price before resolution39Unmarkable on every prior session: no two-sided book and no last trade.
Moved to the benchmark cohort8Graded contracts that are constituents of a governed benchmark. Scored in their own cohort so no observation is counted under two headlines.

436 scored + 15 void + 0 left-censored + 39 unpriced = 490 closed. Of the scored rows, 8 sit in the benchmark cohort and 428 in the cohort above.

Governed benchmarks: scored separately

Benchmark constituents that have been observed resolving are scored in their own cohort, on the same engine and the same sample floors as everything above.

BenchmarkDaysConstituentsOn tapeResolved
WETGRI
5 of 218 tracked constituents have been observed resolving and are scored in the benchmark cohort.
46218165
WETFED
Built by constant-maturity interpolation across Fed meeting strips — it holds no basket, so there are no constituents to grade. Scoring it needs a different instrument (realised policy decisions against the priced path) that is not built.
24n/an/an/a
WETX
3 of 70 tracked constituents have been observed resolving and are scored in the benchmark cohort.
3670183
WETFRAG
318 constituents tracked across 38 published days; 1 is on the price tape and none has resolved yet.
3831810

Per-benchmark counts. The benchmarks share constituents, so the column sums above the 605 distinct contracts tracked overall.

Benchmark Brier score
not published
n = 8 · needs 30

The same metric, over benchmark constituents only. It publishes the moment the cohort clears its floor and not one observation sooner.

Benchmark skill vs base rate
not published
n = 8 · needs 30

Withheld on the same rule. An empty cohort is published as empty rather than borrowing the slate sample's number.

Data we found and refused to score

Prediction snapshot corpus485 graded · excluded (priced-at-or-near-resolution)

485 graded outcomes across 113 settled snapshot files (of 3336 in the corpus), but 59% of them were priced below 5% or above 95% at the time the snapshot was taken — at or after the point the question had effectively resolved. Scoring them would produce a very low Brier that measures the recording date, not the pricing. Excluded until forward-dated snapshots accrue.

What this does and does not measure

It does measure
  • Whether the 428 resolved slate-basket contract prices matched the frequencies that followed.
  • Whether those prices beat knowing only the base rate.
  • Where in the probability range that pricing is strongest and weakest — once the buckets fill.
  • How much data the claim rests on, on every figure, including the figures we withhold.
  • How far the governed benchmarks are from having a track record, in counts rather than adjectives.
It does not measure
  • Whether any index level was “right”. A level has no outcome to be right about.
  • Benchmark calibration beyond the cohort shown above the cohort is young.
  • Anything outside sports, climate, companies, elections, commodities, crypto, economics, culture, politics, tech-science, financials. A category with no resolved contract is absent from this page, not scored as zero.
  • Any return, profit, or what a position would have been worth. Nothing here is a performance claim.
  • W.E.T. forecasting skill. We publish the market’s prices; the market’s calibration is what is scored.
  • Calibration at horizons longer than the tape has observed — at most 10 days so far.

Metrics are withheld below n = 30 for the cohort and n = 30 per bucket. Withholding is implemented in the scoring engine, not in this page, so a withheld number cannot be printed by editing the template. The same object this page renders is served at /api/indices/scorecard — including the nulls and the drop counts, so the arithmetic can be recomputed independently.

Calibration of the contract prices carried on W.E.T.’s daily slate-basket mark tape, measured against the outcomes those venues later reported. It is not a performance claim, not a return, not a forecast and not advice. Governed-benchmark constituents are scored as a separate cohort of 8 observations under the same sample floor; no figure here pools the two. A benchmark level has no outcome and is never scored anywhere.

scorecard/v1 · updated 2026-09-23