check_agent_quality
shallowio.github.fredericmagnathy-ops/alpnai · Verify this server
Check whether a candidate agent workflow regresses before replacing the baseline. Supply baseline and candidate attempts on matching task IDs with client-provided success labels and costs. Returns observed success rates, task-set matching, sample-size, success, latency and cost gates, plus a decision such as collect_more_data, quality_regression or candidate_for_controlled_trial. Optional thresholds use config. It evaluates the recorded labels, not the correctness of answers or future performance. Free calculation; requires an active ALPNAI agent key; no payment or automatic deployment.
1 trials · measured 22 days ago
check_agent_quality scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.fredericmagnathy-ops/alpnai, measured 16 Sept 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- open
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-09-16 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/20e66cb4-1c30-471b-a62a-3ad5367fe77d)