fcastevalbygroup

shallow

io.weathersight/weathersight · Verify this server

Scores how well each forecast model actually performed at each of many locations, by comparing what the models predicted against what was then observed. Returns one score matrix per location, models by evaluation metrics, for a single weather metric at a single lead time, plus the winning model per location. <br><b>When to use:</b> Answer 'which forecast model is most accurate here?' at scale — to map the best model by region, to check whether a specific model is worth using over the default blend, or to compare physical and AI models across the world. <br><b>Date format:</b> ymd_start and ymd_end are YYYYMMDD and inclusive. Inclusive day range as YYYYMMDD. Defaults to the 14 days ending yesterday; today is never included because it has not been verified yet. At most 180 days, which is also how much history the index retains. <br><b>Performance:</b> Filter by ctryid, latlon+radius_km or a small sample_pct to keep it fast; a global 5 percent sample over 14 days is a few hundred locations. <br><b>Prerequisites:</b> None. Supply locid, name, latlon+radius_km, ctryid, or nothing at all with a sample_pct for an unbiased global sample. <br><b>Investigate:</b> Use this to find where a model is strong or weak. Ask for one metric and one lead_time at a time; compare models by reading across each row of evals, and use winners for the answer at a glance. <br><b>Augment:</b> Back a claim about forecast reliability with the measured error of the named model at that place over the recent past. <br><b>Notes:</b> Forecasts come from Open-Meteo. Lead time is the number of days ahead the forecast was issued, so lead_time 1 is yesterday's forecast for today; only the indexed lead times are available. The WN2 model is an ensemble mean, which is smoother than the deterministic models and so tends to score slightly better on mae and rmse and slightly worse on extremes. A score is null, never zero, when its sample was too small to measure: fewer than 5 verified days for mae, rmse, mse, bias and ets, or fewer than 10 for acc. Returns: metric, unit, lead_time, ymd_start, ymd_end, days, models, eval_metrics, sample_pct, sample_pct_effective, seed, buckets, count, scanned, truncated, note, error, results, locid, location, evals, n, winners.

100.0/100

1 trials · measured 22 days ago

fcastevalbygroup scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.weathersight/weathersight, measured 16 Sept 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
not-probed
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-09-16100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: fcastevalbygroup
[![Vouch score](https://vouch.tools/api/tools/3dc7cab1-4e93-44ac-a8bc-b98a08284590/badge.svg)](https://vouch.tools/tools/3dc7cab1-4e93-44ac-a8bc-b98a08284590)
fcastevalbygroup — Vouch