mirror of
https://github.com/internetarchive/heritrix3.git
synced 2026-10-10 06:11:41 +00:00
The paged log viewer gains a search box taking a search-engine style query, passed as the q parameter: status:4xx host:example.com -type:image size>1MB duration>10s - Terms are ANDed; '-' negates; double quotes allow spaces. - 'field:value' is a case-insensitive substring match with '*' wildcards, 'field=value' is exact and 'field~regex' is a Java regex. Numeric fields take the same values with ':' or '=', and no regex; a comparison needs no operator, as in size>1MB. - Bare words and the 'line' field search the whole line in any log. crawl.log also supports status (404, 4xx, <0, 500..599, 404,410), size (>1MB, 100KB..10MB; units are powers of 1024), duration (>10s, 500..2000 in ms), depth (hops from the seed: 0, >5), hops (the hop path, e.g. hops:*E), url, host (includes subdomains), via and viahost (the URI it was discovered from, and its host), type and annot. - Bare words and line terms are highlighted in the results. Filtering keeps the viewer's stateless, byte-position paging and works across checkpoint logs: each request scans forward or backward from pos until it has enough matches or reaches the start or end, streaming a progress bar to the browser. A scan stops after 5 minutes, offering a link to continue from where it stopped, or as soon as the client disconnects. Regex matching is aborted at the same deadline to guard against catastrophic backtracking. The query is carried through all navigation links. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>