judge-batch
shallowio.github.Elimuaa/money-mind · Verify this server
Triage up to 1,000 backtests in one call: the best-of-N multiple-testing selection correction over every return series — $0.01 per series, survivors ranked best-first. Use when your agent has a FARM of backtests, strategies, or research results — dozens to a thousand — and needs triage, not one-by-one verdicts. One call runs the selection correction over every series: each gets NO (killed), PROVISIONAL (survived this gate), or ERROR (malformed). Results ranked sur PAID: $0.01 USDC on Base via x402. Call it to receive the payment challenge. Example request: {"series": [{"returns": [0.8, -0.3, 1.2, 0.5, -0.6, 0.9, 0.4, -0.2, 1.1, 0.7, -0.4, 0.6, 0.3, -0.5, 1.0, 0.8, -0.1, 0.5, 0.9, -0.3], "n_tested": 20, "label": "weak"}, {"returns": [2.0, 2.1, 1.9, 2.2, 2.0, 2.1, 1.8, 2.0,
1 trials · measured 1 day ago
judge-batch scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.Elimuaa/money-mind, measured 6 Oct 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- self-provisionable
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-10-06 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/d4c19866-12fb-40fd-b3df-4376437e4847)