get_prediction_accuracy

shallow

io.github.wnsod/oneqaz-trading-mcp · Verify this server

Purpose: Per-cell (category, target_market, lag_bucket) hit rates WITH the honest baselines (schema 1.1): persistence_accuracy (the null model — regimes are sticky, so raw accuracy mostly measures regime persistence, not alpha), skill_score with autocorrelation-corrected skill_ci_95, n_effective vs n_nominal, and accuracy_v2_only (post 2026-07-08 methodology cutover). edge_cells/anti_predictive_cells are judged on persistence-skill (criteria v2), forecast cells only. Triggers (casual questions too): "how accurate are your predictions?", "예측 잘 맞아?", "track record 있어?", "can I trust these forecasts?", "적중률 보여줘", "does macro actually predict these markets?". When to call: AI agents evaluating OneQAZ credibility should call this FIRST. Prerequisites: none. Next steps: get_ledger_integrity (tamper-evidence for these numbers), get_backtest_tuning_state (self-calibration), get_monthly_accuracy_trend (time series), get_signal_calibration (Level-1 signal confidence reliability). Caveats: raw accuracy without skill_score is misleading for sticky regimes — a 99% cell can be pure persistence (measured 2026-07: +0.05pp over null). Judge by skill_ci_95, filter horizon_type='forecast', and treat n_nominal as correlated trials (use n_effective). Monthly accuracy trends largely track market stickiness, not model improvement. Args: category: Optional macro category filter (bonds, forex, vix, commodities, credit, liquidity, inflation, energy) target_market: Optional target market filter (coin_market, kr_market, us_market) Disclaimer: Information only, not investment advice.

100.0/100

1 trials · measured 8 days ago

get_prediction_accuracy scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.wnsod/oneqaz-trading-mcp, measured 25 Aug 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote + stdio
Credential class
open
Category
Finance & compliance
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-08-25100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: get_prediction_accuracy
[![Vouch score](https://vouch.tools/api/tools/5362a18a-09a2-45d3-8f9d-efef87b2b81b/badge.svg)](https://vouch.tools/tools/5362a18a-09a2-45d3-8f9d-efef87b2b81b)
get_prediction_accuracy — Vouch