assay_conformal

shallow

com.alphaassay/mcp · Verify this server

Use this when your model emits prediction intervals or confidence bands and you want to audit whether realised outcomes actually fall inside them at the claimed coverage -- a calibration audit, not advice. Does your model's confidence label survive contact with outcomes? Coverage audit over prediction intervals -- confidence claims are the third claim type after return claims and risk forecasts. Submit the prediction intervals your ML model produced (lower and upper bounds, one pair per point), the realised outcomes, and the coverage the model claims (e.g. 0.9). The miss count is graded against the EXACT distribution that claim implies: binomial for a pointwise claim, or -- when you disclose the split-conformal calibration_size -- the beta-binomial the split-conformal guarantee actually promises (Vovk 2012: realised coverage is Beta-distributed around the claim), so a CORRECT procedure with a small calibration set is never punished for its honest variance. Zones reuse the same traffic-light boundaries as the VaR cell (green below cumulative probability 0.95, yellow to 0.9999, red above); red earns the named demote CONFORMAL_COVERAGE_SHORTFALL. Efficiency is disclosed, never assumed: intervals wider than the whole observed outcome range flag the advisory conformal_intervals_uninformative (coverage bought by construction), over-conservative intervals flag conformal_overcovered -- advisories, never kills, demote-only throughout. Code-computed end to end, fail-closed on malformed, inverted, undersized or oversized input. References: Vovk/Gammerman/Shafer 2005; Lei et al., JASA 2018. NOT financial advice; no order path. Price: per check; see https://api.alphaassay.com/v1/meta/pricing (api_key required -- account setup at https://api.alphaassay.com/account).

100.0/100

1 trials · measured 8 days ago

assay_conformal scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against com.alphaassay/mcp, measured 25 Aug 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
self-provisionable
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-08-25100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: assay_conformal
[![Vouch score](https://vouch.tools/api/tools/b203e5a2-ed4b-474f-b569-ed2938beaf26/badge.svg)](https://vouch.tools/tools/b203e5a2-ed4b-474f-b569-ed2938beaf26)
assay_conformal — Vouch