get_calibration

shallow

com.anteproof/anteproof · Verify this server

The calibration record: Brier skill by question family for the pastcast pool and the live track, with the launch-gate verdict for the horizon asked for. horizon: 7, 30, 90, long, or all (default: the horizon the latest verdict scopes on); the 90-day maintenance families (commits_90d, release_90d) are at 90. regime.gate is the verdict evaluated on THAT horizon, or null when no gate has been run on it — never another horizon's verdict, so null means those families are unjudged, not that they failed. regime.latest_gate is the newest verdict overall whatever its horizon, carrying its own domain and horizon_days; read applies_to_this_view before treating it as a verdict on the families returned. SHAPE: `pastcast` and `live` each carry the track's own statistics plus `families`, a list of per-family blocks (family, label, n, brier, brier_skill_score, base_rate, mean_p, z_bias, hit_rate, bins, enough_to_judge). Every one of those blocks, and each track, also carries a `withheld` sibling: null when the numbers are real, otherwise an object {reason, metrics, degraded_from, detail}. WHEN `withheld` IS NON-NULL, EVERY NUMBER IN THAT BLOCK IS JSON null — n included. null there means 'we refuse to publish this', NOT zero and NOT 'nothing resolved'; do not coerce it to 0, do not average over it, and do not report it as a score. A real n=0 with `withheld: null` is the honest 'nothing has resolved yet'. The families over GH Archive activity series — commits_90d, release_90d and the legacy daily families commits_hourly, release_published and stars_hourly — are withheld unconditionally today, with reason 'resolutions_untrusted': every one of those questions was resolved by counting a GH Archive series whose capture collapsed, so no score over that pool is sound at any horizon. `detail` is a sentence you can quote to the user. Other reasons a number can be withheld are 'source_degraded', 'source_collapsed' and 'insufficient_data'. Every other family (news tone, Wikipedia pageviews, Hacker News points, job postings, crates and PyPI downloads, prediction-market questions) publishes real numbers as before.

100.0/100

1 trials · measured 27 days ago

get_calibration scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against com.anteproof/anteproof, measured 11 Sept 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
gated
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-09-11100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: get_calibration
[![Vouch score](https://vouch.tools/api/tools/aef2a4a7-0fd0-402e-aa80-8c1e0b59057f/badge.svg)](https://vouch.tools/tools/aef2a4a7-0fd0-402e-aa80-8c1e0b59057f)
get_calibration — Vouch