fcastevalbygroup
shallowio.weathersight/weathersight · Verify this server
Scores how well each forecast model actually performed at each of many locations, by comparing what the models predicted against what was then observed. Returns one score matrix per location, models by evaluation metrics, for a single weather metric at a single lead time, plus the winning model per location. <br><b>When to use:</b> Answer 'which forecast model is most accurate here?' at scale — to map the best model by region, to check whether a specific model is worth using over the default blend, or to compare physical and AI models across the world. <br><b>Date format:</b> ymd_start and ymd_end are YYYYMMDD and inclusive. Inclusive day range as YYYYMMDD. Defaults to the 14 days ending yesterday; today is never included because it has not been verified yet. At most 180 days, which is also how much history the index retains. <br><b>Performance:</b> Filter by ctryid, latlon+radius_km or a small sample_pct to keep it fast; a global 5 percent sample over 14 days is a few hundred locations. <br><b>Prerequisites:</b> None. Supply locid, name, latlon+radius_km, ctryid, or nothing at all with a sample_pct for an unbiased global sample. <br><b>Investigate:</b> Use this to find where a model is strong or weak. Ask for one metric and one lead_time at a time; compare models by reading across each row of evals, and use winners for the answer at a glance. <br><b>Augment:</b> Back a claim about forecast reliability with the measured error of the named model at that place over the recent past. <br><b>Notes:</b> Forecasts come from Open-Meteo. Lead time is the number of days ahead the forecast was issued, so lead_time 1 is yesterday's forecast for today; only the indexed lead times are available. The WN2 model is an ensemble mean, which is smoother than the deterministic models and so tends to score slightly better on mae and rmse and slightly worse on extremes. A score is null, never zero, when its sample was too small to measure: fewer than 5 verified days for mae, rmse, mse, bias and ets, or fewer than 10 for acc. Returns: metric, unit, lead_time, ymd_start, ymd_end, days, models, eval_metrics, sample_pct, sample_pct_effective, seed, buckets, count, scanned, truncated, note, error, results, locid, location, evals, n, winners.
1 trials · measured 22 days ago
fcastevalbygroup scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.weathersight/weathersight, measured 16 Sept 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- not-probed
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-09-16 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/3dc7cab1-4e93-44ac-a8bc-b98a08284590)