check_eval_selection
shallowio.github.Elimuaa/money-mind · Verify this server
FREE. Use after picking the best of several prompts, models, configs or hyperparameters. Keeping the top scorer out of N is SELECTION: the winner's score is inflated simply by having looked N times, and every eval harness reports it as though one experiment were run. Give the score each variant achieved and it returns how much of the winner's margin the search itself explains, what pure noise would have handed you, and whether the excess is real. Supply n_per_variant to also learn whether your variants can be told apart at all. Call this before concluding an optimisation found something.
1 trials · measured 1 day ago
check_eval_selection scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.Elimuaa/money-mind, measured 6 Oct 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- self-provisionable
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-10-06 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/89ae2141-505f-44fd-bf75-3a8ebeb67d5e)