Every contract W.E.T. carries is a binary question that eventually resolves YES or NO at its venue, so the prices can be scored against what actually happened. This page reports exactly one thing: the calibration of the 140 contract prices the daily slate-basket mark tape has so far seen priced while open and later observed resolving. Which contracts those are, and which they are not, is stated before any number.
It is the slate-basket mark tape. The daily slate-basket mark tape — the two-sided price recorder that marks every holding in every published W.E.T. slate basket, once per session. Across 6 sessions (2026-08-01 → 2026-08-06) it covers 26 of the 26 slate books on record — counted by checking each book’s holdings against the tape, not by counting books — and 140 of those holdings have resolved.
It is not the governed benchmarks — WETGRI, WETFED, WETX and WETFRAG. Not one of these 140 observations is a benchmark constituent. The first benchmark settlement row was published 2026-08-06, and the longest-running ledger holds 1 published day; of the 140 constituents it tracks, none has resolved yet. Benchmark calibration is therefore not measurable, and this page reports that rather than estimating it from a sample of something else.
And no index level is ever scored. Nothing that happens later makes a printed level right or wrong — an index reads prices, it does not forecast anything. That holds permanently, for benchmarks and slates alike, however large this sample grows.
Read off the category frozen on each holding when the basket opened it, not inferred here.
A contract can sit in several baskets, so these are memberships rather than a partition and can sum above 140.
Every contract held in a published W.E.T. SLATE basket that the daily mark tape saw priced while the question was open and later observed resolving at its venue. Slate baskets only — it excludes the governed benchmarks (WETGRI, WETFED, WETX, WETFRAG), whose constituents are covered by the separate benchmark cohort.
Every figure below describes those 140 contracts and nothing else — not the benchmarks, not the desk as a whole, and not any position or result.
Mean squared error of the price against the outcome. Lower is better. A constant 50% forecast scores 0.25.
Against a forecaster who knew only how often these contracts resolve YES. Positive means the prices carried real information; a raw Brier alone cannot tell you this.
Share of the sample that resolved YES. The reference forecast the skill score is measured against.
Secondary metric, clamped to stay finite. It punishes confident misses far harder than the Brier, which makes it noisy on small samples.
The table a single score cannot fake: a systematic bias hides inside an average but not inside ten rows. A bucket publishes its observed frequency only once it holds 30 graded outcomes. Below that the value is withheld and the row says how many more it needs — an observed frequency built from six contracts is ±20 points of binomial noise, which is larger than any miscalibration worth reporting.
| Priced at | Contracts | Mean priced | Resolved YES | Observed freq. | Gap |
|---|---|---|---|---|---|
0–10% | 17 | 2.2% | 0 | needs 13 more | — |
10–20% | 4 | 12.8% | 1 | needs 26 more | — |
20–30% | 2 | 23.5% | 0 | needs 28 more | — |
30–40% | 8 | 37.3% | 3 | needs 22 more | — |
40–50% | 14 | 45.1% | 6 | needs 16 more | — |
50–60% | 33 | 54.2% | 21 | 63.6% | +9.5pp |
60–70% | 18 | 62.8% | 11 | needs 12 more | — |
70–80% | 9 | 73.1% | 7 | needs 21 more | — |
80–90% | 11 | 86.5% | 10 | needs 19 more | — |
90–100% | 24 | 97.0% | 24 | needs 6 more | — |
The scored count is a subset of everything the tape saw close. Publishing the rejects is the only way to show the sample was not selected after the fact — so rows 2–5 partition row 1 exactly, and the last row splits the scored contracts between the two cohorts on this page.
| Contracts the tape saw close | 187 | The denominator. Distinct contracts that reached any terminal state — resolved YES, resolved NO, or voided. |
| Scored | 140 | A live price exists on an earlier tape day than the day the answer was observed. |
| Dropped — voided | 8 | Closed by the venue with no gradeable answer. Scoring a void as NO would punish a price for a question never asked. |
| Dropped — already resolved when first seen | 0 | We never observed a live price. Pairing the settlement mark with itself would drive the Brier to zero — the classic way a fake scorecard looks superb. |
| Dropped — no live price before resolution | 39 | Unmarkable on every prior session: no two-sided book and no last trade. |
| Moved to the benchmark cohort | 0 | Graded contracts that are constituents of a governed benchmark. Scored in their own cohort so no observation is counted under two headlines. |
140 scored + 8 void + 0 left-censored + 39 unpriced = 187 closed. Of the scored rows, 0 sit in the benchmark cohort and 140 in the cohort above.
The benchmarks are the governed class, so they are the ones a reader most wants a track record for. They do not have one yet, and the useful form of that answer is a count rather than a caption. The longest-running of the 4 benchmark ledgers holds 1 published day; together they track 140 distinct constituents, of which 0 have resolved. The cohort is computed live on every render, so it fills itself in the day that changes — no one has to remember to switch it on.
| Benchmark | Days | Constituents | On tape | Resolved |
|---|---|---|---|---|
WETGRI 80 constituents tracked across 1 published day; 12 are on the price tape and none has resolved yet. | 1 | 80 | 12 | 0 |
WETFED Built by constant-maturity interpolation across Fed meeting strips — it holds no basket, so there are no constituents to grade. Scoring it needs a different instrument (realised policy decisions against the priced path) that is not built. | 1 | n/a | n/a | n/a |
WETX 40 constituents tracked across 1 published day; 8 are on the price tape and none has resolved yet. | 1 | 40 | 8 | 0 |
WETFRAG 21 constituents tracked across 1 published day; 1 is on the price tape and none has resolved yet. | 1 | 21 | 1 | 0 |
Per-benchmark counts. The benchmarks share constituents, so the column sums above the 140 distinct contracts tracked overall.
The same metric, over benchmark constituents only. It publishes the moment the cohort clears its floor and not one observation sooner.
Withheld on the same rule. An empty cohort is published as empty rather than borrowing the slate sample's number.
125 graded outcomes across 19 settled snapshot files (of 1307 in the corpus), but 81% of them were priced below 5% or above 95% at the time the snapshot was taken — at or after the point the question had effectively resolved. Scoring them would produce a very low Brier that measures the recording date, not the pricing. Excluded until forward-dated snapshots accrue.
Metrics are withheld below n = 30 for the cohort and n = 30 per bucket. Withholding is implemented in the scoring engine, not in this page, so a withheld number cannot be printed by editing the template. The same object this page renders is served at /api/indices/scorecard — including the nulls and the drop counts, so the arithmetic can be recomputed independently.
Calibration of the contract prices carried on W.E.T.’s daily slate-basket mark tape, measured against the outcomes those venues later reported. It is not a performance claim, not a return, not a forecast and not advice. It does NOT cover the governed benchmarks: none of the 140 tracked benchmark constituents has been observed resolving, so benchmark calibration is reported as unmeasurable rather than estimated. A benchmark level has no outcome and is never scored anywhere.
scorecard/v1 · updated 2026-08-06