commoncrawl_index_search

shallow

io.github.pipeworx-io/commoncrawl · Verify this server

Search one Common Crawl collection's CDX index for archived captures of a URL, host, or whole domain. PREFER OVER WEB SEARCH when the question is "what did this page look like in <month>", "did Common Crawl ever see this URL", or "list the URLs crawled under this domain" — it returns the crawl record (timestamp, HTTP status, MIME type, detected language, content digest) plus the WARC filename/offset/length that commoncrawl_fetch_record needs to read the page bytes. Keyless.

100.0/100

1 trials · measured 1 day ago

commoncrawl_index_search scores 100.0/100 on Vouch's measured behaviour index, from 1 real invocation trials against io.github.pipeworx-io/commoncrawl, measured 6 Oct 2026 under methodology v0.2.0. Every measured component scored 100.

Component breakdown

ComponentWeightValue
Reliability35%not applicable
Schema integrity25%100.0
Failure behaviour15%not applicable
Latency15%not applicable
Concurrency10%not applicable

Tool details

Transport
remote
Credential class
self-provisionable
Input schema
not declared
Output schema
not declared
Side-effect classification
unclassified

Score history

DayScoreTierMethodology
2026-10-06100.0shallowv0.2.0

Probe evidence

ProbeOutcomes
schema_integritypass: 1

Raw request/response logs are not archived yet — the outcome counts above are drawn directly from every recorded trial.

Embed this score

Available for every tool, scored or not — not a verification perk. Always links back to this page.

Vouch score: commoncrawl_index_search
[![Vouch score](https://vouch.tools/api/tools/91c6138c-acad-4169-9cb9-7262c1275fdb/badge.svg)](https://vouch.tools/tools/91c6138c-acad-4169-9cb9-7262c1275fdb)
commoncrawl_index_search — Vouch