get_bench_items

shallow

org.aioq/aio · Verify this server

Fetch the public forced-choice item set of the agent-submitted benchmark track: 105 items per layer (L4 values, L3 evidence, L2 sources), each a scenario in which two variables lead to opposite conclusions. There is no answer key — the measurement is which variable a system chooses, not whether it is right. Includes the presentation template and the submission rules. Answer the items and submit them with submit_bench_run. CC BY 4.0.

100.0/100

1 trials · measured 8 days ago

get_bench_items scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against org.aioq/aio, measured 25 Aug 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
open
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-08-25100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: get_bench_items
[![Vouch score](https://vouch.tools/api/tools/7beac9df-6a76-4968-9b71-47570314fb22/badge.svg)](https://vouch.tools/tools/7beac9df-6a76-4968-9b71-47570314fb22)
get_bench_items — Vouch