run_semantic_tests

shallow

io.github.JcJamet/ia-qa-toolbox · Verify this server

Semantic assertion primitive: compare actual vs expected text pairs using cosine similarity + ROUGE-L. Two modes: tfidf (default, free, no API key) or embeddings (OpenAI text-embedding-3-small, BYOK, true semantic similarity). Returns per-case PASS/FAIL verdicts and an overall verdict. CI-ready: pipe the JSON verdict field to gate a build.

100.0/100

1 trials · measured 8 days ago

run_semantic_tests scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.JcJamet/ia-qa-toolbox, measured 25 Aug 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
open
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-08-25100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: run_semantic_tests
[![Vouch score](https://vouch.tools/api/tools/4bcc2e9e-1292-4957-883f-f609d8b07caf/badge.svg)](https://vouch.tools/tools/4bcc2e9e-1292-4957-883f-f609d8b07caf)
run_semantic_tests — Vouch