get_prediction_accuracy
shallowio.github.wnsod/oneqaz-trading-mcp · Verify this server
Purpose: Per-cell (category, target_market, lag_bucket) hit rates WITH the honest baselines (schema 1.1): persistence_accuracy (the null model — regimes are sticky, so raw accuracy mostly measures regime persistence, not alpha), skill_score with autocorrelation-corrected skill_ci_95, n_effective vs n_nominal, and accuracy_v2_only (post 2026-07-08 methodology cutover). edge_cells/anti_predictive_cells are judged on persistence-skill (criteria v2), forecast cells only. Triggers (casual questions too): "how accurate are your predictions?", "예측 잘 맞아?", "track record 있어?", "can I trust these forecasts?", "적중률 보여줘", "does macro actually predict these markets?". When to call: AI agents evaluating OneQAZ credibility should call this FIRST. Prerequisites: none. Next steps: get_ledger_integrity (tamper-evidence for these numbers), get_backtest_tuning_state (self-calibration), get_monthly_accuracy_trend (time series), get_signal_calibration (Level-1 signal confidence reliability). Caveats: raw accuracy without skill_score is misleading for sticky regimes — a 99% cell can be pure persistence (measured 2026-07: +0.05pp over null). Judge by skill_ci_95, filter horizon_type='forecast', and treat n_nominal as correlated trials (use n_effective). Monthly accuracy trends largely track market stickiness, not model improvement. Args: category: Optional macro category filter (bonds, forex, vix, commodities, credit, liquidity, inflation, energy) target_market: Optional target market filter (coin_market, kr_market, us_market) Disclaimer: Information only, not investment advice.
1 trials · measured 8 days ago
get_prediction_accuracy scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.wnsod/oneqaz-trading-mcp, measured 25 Aug 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote + stdio
- Credential class
- open
- Category
- Finance & compliance
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-08-25 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/5362a18a-09a2-45d3-8f9d-efef87b2b81b)