fcastevalbylocation
shallowio.weathersight/weathersight · Verify this server
Scores how well each forecast model performed at ONE location, across several forecast lead times and across a sliding time window. Shows both how fast accuracy decays as the forecast reaches further ahead, and how accuracy at a given lead time has varied over recent weeks. <br><b>When to use:</b> Answer 'how far ahead can I trust the forecast here?' or 'has the forecast been unusually poor here lately?' Use fcastevalbygroup instead to compare many locations at one lead time. <br><b>Date format:</b> ymd_start and ymd_end are YYYYMMDD and inclusive. Inclusive day range as YYYYMMDD. Defaults to the 14 days ending yesterday; today is never included because it has not been verified yet. At most 180 days, which is also how much history the index retains. <br><b>Performance:</b> Fast: one location. The cost grows with the number of lead times and the length of the range. <br><b>Prerequisites:</b> A location: locid from /api/location, or name, or latlon. <br><b>Investigate:</b> Ask for several lead_times at once (e.g. 1,3,5,7) to get the skill-decay curve in one call. Set window to get a moving average and see the trend over time rather than one number. <br><b>Augment:</b> Quantify how reliable a forecast for this place actually is at the lead time being written about. <br><b>Notes:</b> Forecasts come from Open-Meteo. Lead time is the number of days ahead the forecast was issued; only the indexed lead times are available. A score is null, never zero, when its sample was too small to measure: fewer than 5 verified days for mae, rmse, mse, bias and ets, or fewer than 10 for acc. Returns: locid, location, metric, unit, ymd_start, ymd_end, days, window, num_windows, windows, lead_times, models, eval_metrics, scanned, count, truncated, error, evals, n.
1 trials · measured 22 days ago
fcastevalbylocation scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.weathersight/weathersight, measured 16 Sept 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- not-probed
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-09-16 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/c1706656-9cb3-478d-81a3-f821e5a21fb5)