Files
heritrix3/engine
Alex OsborneandClaude Opus 5.5 f5070f6203 Paged log viewer: search and filter with a small query language
The paged log viewer gains a search box taking a search-engine style
query, passed as the q parameter:

  status:4xx host:example.com -type:image size>1MB duration>10s

- Terms are ANDed; '-' negates; double quotes allow spaces.
- 'field:value' is a case-insensitive substring match with '*'
  wildcards, 'field=value' is exact and 'field~regex' is a Java regex.
  Numeric fields take the same values with ':' or '=', and no regex;
  a comparison needs no operator, as in size>1MB.
- Bare words and the 'line' field search the whole line in any log.
  crawl.log also supports status (404, 4xx, <0, 500..599, 404,410),
  size (>1MB, 100KB..10MB; units are powers of 1024), duration (>10s,
  500..2000 in ms), depth (hops from the seed: 0, >5), hops (the hop
  path, e.g. hops:*E), url, host (includes subdomains), via and viahost
  (the URI it was discovered from, and its host), type and annot.
- Bare words and line terms are highlighted in the results.

Filtering keeps the viewer's stateless, byte-position paging and works
across checkpoint logs: each request scans forward or backward from pos
until it has enough matches or reaches the start or end, streaming a
progress bar to the browser. A scan stops after 5 minutes, offering a
link to continue from where it stopped, or as soon as the client
disconnects. Regex matching is aborted at the same deadline to guard against
catastrophic backtracking. The query is carried through all navigation links.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 16:54:19 +09:00
..
2009-05-11 22:56:36 +00:00