find_evidence
shallowsite.chatgpt.tuned-drake-1114.execution-evidence-lab/agent-workspace-notes · Verify this server
Use when debugging a public Python error and looking for an existing measured reproduction. Search by exact error or library name alone (library browsing). Longer problems need a measured-topic clue; a lone library mention does not imply a matching fix. Curated examples: asyncio, gather, TaskGroup, cancellation, ExceptionGroup, numpy.dtype size changed, binary incompatibility, NumPy 2, pandas, PyArrow, pydantic partial update, pydantic PATCH, exclude_unset, exclude_none, model_fields_set, pydantic-settings, pydantic, extra_forbidden, Extra inputs are not permitted, dotenv, SQLAlchemy, AsyncSession, AsyncAttrs, MissingGreenlet, greenlet_spawn has not been called, sqlite, sqlalchemy, no such table, in memory, StaticPool, sqlite3, savepoint, rollback, release, httpx, starlette, TestClient, unexpected keyword argument app, lifespan, subprocess, Popen, PIPE, communicate, TimeoutExpired, urllib.parse, urljoin, same origin, URL prefix, userinfo, datetime, zoneinfo, DST, elapsed time, fold. Returns ranked previews with record_id, test_id, scope, platform, page_url and separately scoped supplementary_evidence when available; no primary match returns an empty records array. Separately typed research_supplements may provide existing non-Python research files, e.g. MCP cancellation/retry accounting in SDK v1.30.0; inspect their version limits and file URLs, not read_evidence. A supplement is not a primary record or upstream resolution. Read the preview before choosing read_evidence for receipt-free retrieval; get_evidence additionally creates an optional receipt and private report proof. Full records and files are freely readable. Not a general web search or proof of compatibility. Requests are logged.
1 trials · measured 2 days ago
find_evidence scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against site.chatgpt.tuned-drake-1114.execution-evidence-lab/agent-workspace-notes, measured 6 Oct 2026 under methodology v0.2.0. Every measured component scored 100.
Component breakdown
| Component | Weight | Value |
|---|---|---|
| Reliability | 35% | not applicable |
| Schema integrity | 25% | 100.0 |
| Failure behaviour | 15% | not applicable |
| Latency | 15% | not applicable |
| Concurrency | 10% | not applicable |
Tool details
- Transport
- remote
- Credential class
- self-provisionable
- Input schema
- not declared
- Output schema
- not declared
- Side-effect classification
- unclassified
Score history
| Day | Score | Tier | Methodology |
|---|---|---|---|
| 2026-10-06 | 100.0 | shallow | v0.2.0 |
Probe evidence
| Probe | Outcomes |
|---|---|
| schema_integrity | pass: 1 |
Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.
Embed this score
Available for every tool, scored or not — not a verification perk. Always links back to this page.
[](https://vouch.tools/tools/688f9490-d562-4c42-8c2a-77b8150224b9)