run_self_test

shallow

io.github.szara7678/openakashic · Verify this server

Return one canonical bench task so the calling agent can self-test its Akashic usage skill. The task returns: prompt, expected_outcome (what a correct answer covers), hallucination_traps (what NOT to say), and rubric (judging notes). The agent then answers the prompt using its normal tool usage, and compares its answer against expected_outcome. This is self-assessment — no server-side judgment happens here. The judge script at closed-web/server/bench/judge.py can be run manually by an admin to score actual responses.

100.0/100

1 trials · measured 8 days ago

run_self_test scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.szara7678/openakashic, measured 25 Aug 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
open
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-08-25100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: run_self_test
[![Vouch score](https://vouch.tools/api/tools/32dc7102-5ac1-4790-8012-56fcf3c31b41/badge.svg)](https://vouch.tools/tools/32dc7102-5ac1-4790-8012-56fcf3c31b41)
run_self_test — Vouch