Restlet's Directory gives every file an Expires header 10 minutes in the
future. Job files (logs, reports, crawler-beans.cxml) change constantly, so
browsers could show stale copies. Worst case, the config editor loads a
cached crawler-beans.cxml and saving it silently reverts newer changes.
Clear the expiration date and send Cache-Control: no-cache instead, so
browsers revalidate against Last-Modified on each use.
Add CrossSiteRequestFilter in front of the web UI's router to protect
against cross-site request forgery. POST, PUT and DELETE requests are
rejected with 403 when the browser's Sec-Fetch-Site header says they
came from another site.
Sec-Fetch-Site is set by the browser from its own view of the page and
target URL, so unlike comparing Origin with Host it isn't affected by
reverse proxies rewriting the host or scheme. Requests without the
header, such as from curl and other API clients, are allowed.
Origins listed in the heritrix.trustedOrigins system property (comma
separated) are allowed to send cross-site requests.
Only write heritrix_dmesg.log when requested
Heritrix.java always created heritrix_dmesg.log in $HERITRIX_HOME, even
though only the bin/heritrix script reads it, and only when it starts
Heritrix in the background. This got in the way of making
$HERITRIX_HOME read-only for Docker and systemd setups.
The startup log is now written only when the heritrix.dmesg system
property gives its path. bin/heritrix sets this property in background
mode and keeps the file in the same place as before.
Also fix a typo in bin/heritrix ($startmessage) that stopped it
removing a heritrix_dmesg.log left over from a previous start.
Set -Dheritrix.dmesg in heritrix.cmd too.
The paged log viewer gains a search box taking a search-engine style
query, passed as the q parameter:
status:4xx host:example.com -type:image size>1MB duration>10s
- Terms are ANDed; '-' negates; double quotes allow spaces.
- 'field:value' is a case-insensitive substring match with '*'
wildcards, 'field=value' is exact and 'field~regex' is a Java regex.
Numeric fields take the same values with ':' or '=', and no regex;
a comparison needs no operator, as in size>1MB.
- Bare words and the 'line' field search the whole line in any log.
crawl.log also supports status (404, 4xx, <0, 500..599, 404,410),
size (>1MB, 100KB..10MB; units are powers of 1024), duration (>10s,
500..2000 in ms), depth (hops from the seed: 0, >5), hops (the hop
path, e.g. hops:*E), url, host (includes subdomains), via and viahost
(the URI it was discovered from, and its host), type and annot.
- Bare words and line terms are highlighted in the results.
Filtering keeps the viewer's stateless, byte-position paging and works
across checkpoint logs: each request scans forward or backward from pos
until it has enough matches or reaches the start or end, streaming a
progress bar to the browser. A scan stops after 5 minutes, offering a
link to continue from where it stopped, or as soon as the client
disconnects. Regex matching is aborted at the same deadline to guard against
catastrophic backtracking. The query is carried through all navigation links.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Checkpointing renames each log in place (crawl.log becomes
crawl.log.cp00001-<timestamp>) and starts a new one, so earlier entries
could only be viewed one file at a time. The paged viewer now offers an
"include N checkpoint logs" link (all=y) that shows the log and its
rotated generations in the current launch as one virtual file, oldest
first, with each run of lines labelled with its file.
Positions in the virtual file stay valid across checkpoints, since a
rotated file keeps its place and the new log continues after it, and
they are converted when toggling so the view stays where it was.
Gzipped generations can't be read at arbitrary positions and are
left out.
LogSeries opens the files and fixes their lengths up front, so a
rotation mid-request doesn't change what is read. Paging uses a new
FilteredLineScanner over a LogSeries instead of FileUtils.pagedLines,
which can only read a single file. Results are the same for every
position the navigation links generate; a hand-edited pos in the
middle of a line now starts at that line rather than the next one.
Lines are capped at 256 KiB so a file with few or no newlines can't
exhaust the heap.
Also fixes in the crawl.log rendering: a rotated crawl.log gets the
same status highlighting, status codes unknown to Jetty (e.g. 999) and
numbers too long to be a status no longer fail the page, and the file
name, status titles and navigation links are HTML escaped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This fixes a bug where POST requests would get recorded with a doubled header and a bunch of nulls at the end.
It seems that the onRequestContent() listener doesn't pass us the right buffer when proxying. The buffer seems to contain client-to-server data and the buffer's limit isn't set. So instead this records the request content by overriding newProxyToServerRequestContent.
Foundation 4 has been obsolete since 2013, and these libraries trigger security alerts. We're also not using them to do much, just the modal dialogs, alert close buttons and collapsing the top-bar on mobile. This replaces the modals with the native HTML dialog element and implements the trivial click event handlers needed for hiding and showing things.
This means:
- Job profiles can be used without having to hunt for them among the regular jobs.
- The Groovy default profile can be selected as an alternative to Spring XML one.
- The default profiles can be customised by creating a job profile with the same name.
BdbMultipleWorkQueues.calculateInsertKey() silently clips precedence
values to 127 for queue ordering. This can cause unexpected behavior
for custom cost assignment policies that return values above 127.
Log a warning when clipping occurs and fix the misleading comment in
CrawlURI.holderCost that said "should not exceed 255" (the actual
limit is 127, not 255).
Fixes#502
Heritrix's Spring XML based job config format by design allows arbitrary code execution, so this doesn't realistically make things much safer, but I also don't see any reason we need external entity resolution enabled. I suppose there's a chance this could mitigate a generic automated attack that uses XXE but doesn't target Heritrix or Spring XML specifically.
Fixes#711
This adds context-sensitive autocompletion for bean classes, bean ids, property names, basic Spring XML tags and attributes. The bean completions use a new `/engine/beandoc` endpoint that serves the combination of all the `/META-INF/heritrix-beans.json` files generated at compile-time by the heritrix-docgen annotation processor.