find_evidence

shallow

site.chatgpt.tuned-drake-1114.execution-evidence-lab/agent-workspace-notes · Verify this server

Use when debugging a public Python error and looking for an existing measured reproduction. Search by exact error or library name alone (library browsing). Longer problems need a measured-topic clue; a lone library mention does not imply a matching fix. Curated examples: asyncio, gather, TaskGroup, cancellation, ExceptionGroup, numpy.dtype size changed, binary incompatibility, NumPy 2, pandas, PyArrow, pydantic partial update, pydantic PATCH, exclude_unset, exclude_none, model_fields_set, pydantic-settings, pydantic, extra_forbidden, Extra inputs are not permitted, dotenv, SQLAlchemy, AsyncSession, AsyncAttrs, MissingGreenlet, greenlet_spawn has not been called, sqlite, sqlalchemy, no such table, in memory, StaticPool, sqlite3, savepoint, rollback, release, httpx, starlette, TestClient, unexpected keyword argument app, lifespan, subprocess, Popen, PIPE, communicate, TimeoutExpired, urllib.parse, urljoin, same origin, URL prefix, userinfo, datetime, zoneinfo, DST, elapsed time, fold. Returns ranked previews with record_id, test_id, scope, platform, page_url and separately scoped supplementary_evidence when available; no primary match returns an empty records array. Separately typed research_supplements may provide existing non-Python research files, e.g. MCP cancellation/retry accounting in SDK v1.30.0; inspect their version limits and file URLs, not read_evidence. A supplement is not a primary record or upstream resolution. Read the preview before choosing read_evidence for receipt-free retrieval; get_evidence additionally creates an optional receipt and private report proof. Full records and files are freely readable. Not a general web search or proof of compatibility. Requests are logged.

100.0/100

1 trials · measured 2 days ago

find_evidence scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against site.chatgpt.tuned-drake-1114.execution-evidence-lab/agent-workspace-notes, measured 6 Oct 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
self-provisionable
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-10-06100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: find_evidence
[![Vouch score](https://vouch.tools/api/tools/688f9490-d562-4c42-8c2a-77b8150224b9/badge.svg)](https://vouch.tools/tools/688f9490-d562-4c42-8c2a-77b8150224b9)
find_evidence — Vouch