Commit Graph
765 Commits
Author SHA1 Message Date
Alex Osborne d6747686e3 Don't let browsers cache job directory files for 10 minutes (#782)
Restlet's Directory gives every file an Expires header 10 minutes in the
future. Job files (logs, reports, crawler-beans.cxml) change constantly, so
browsers could show stale copies. Worst case, the config editor loads a
cached crawler-beans.cxml and saving it silently reverts newer changes.

Clear the expiration date and send Cache-Control: no-cache instead, so
browsers revalidate against Last-Modified on each use.
2026-10-07 12:42:02 +09:00
Alex Osborne 7939d0d7e1 Reject cross-site requests to the web UI
Add CrossSiteRequestFilter in front of the web UI's router to protect
against cross-site request forgery. POST, PUT and DELETE requests are
rejected with 403 when the browser's Sec-Fetch-Site header says they
came from another site.

Sec-Fetch-Site is set by the browser from its own view of the page and
target URL, so unlike comparing Origin with Host it isn't affected by
reverse proxies rewriting the host or scheme. Requests without the
header, such as from curl and other API clients, are allowed.

Origins listed in the heritrix.trustedOrigins system property (comma
separated) are allowed to send cross-site requests.
2026-10-06 12:21:40 +09:00
Alex Osborne 55bef0aa32 Merge pull request #780 from internetarchive/paged-log-viewer-search
Crawl log search
2026-10-06 12:09:09 +09:00
Alex Osborne 6f9a82c8fb Fix HTML escaping in templates
Enables autoescaping in Freemarker and adds some missing escaping
where we generate HTML in Java methods.
2026-10-06 12:04:23 +09:00
Alex Osborne a09f38570d Only write heritrix_dmesg.log when requested (#779)
Only write heritrix_dmesg.log when requested

Heritrix.java always created heritrix_dmesg.log in $HERITRIX_HOME, even
though only the bin/heritrix script reads it, and only when it starts
Heritrix in the background. This got in the way of making
$HERITRIX_HOME read-only for Docker and systemd setups.

The startup log is now written only when the heritrix.dmesg system
property gives its path. bin/heritrix sets this property in background
mode and keeps the file in the same place as before.

Also fix a typo in bin/heritrix ($startmessage) that stopped it
removing a heritrix_dmesg.log left over from a previous start.

Set -Dheritrix.dmesg in heritrix.cmd too.
2026-10-06 11:29:51 +09:00
Alex OsborneandClaude Opus 5.5 f5070f6203 Paged log viewer: search and filter with a small query language
The paged log viewer gains a search box taking a search-engine style
query, passed as the q parameter:

  status:4xx host:example.com -type:image size>1MB duration>10s

- Terms are ANDed; '-' negates; double quotes allow spaces.
- 'field:value' is a case-insensitive substring match with '*'
  wildcards, 'field=value' is exact and 'field~regex' is a Java regex.
  Numeric fields take the same values with ':' or '=', and no regex;
  a comparison needs no operator, as in size>1MB.
- Bare words and the 'line' field search the whole line in any log.
  crawl.log also supports status (404, 4xx, <0, 500..599, 404,410),
  size (>1MB, 100KB..10MB; units are powers of 1024), duration (>10s,
  500..2000 in ms), depth (hops from the seed: 0, >5), hops (the hop
  path, e.g. hops:*E), url, host (includes subdomains), via and viahost
  (the URI it was discovered from, and its host), type and annot.
- Bare words and line terms are highlighted in the results.

Filtering keeps the viewer's stateless, byte-position paging and works
across checkpoint logs: each request scans forward or backward from pos
until it has enough matches or reaches the start or end, streaming a
progress bar to the browser. A scan stops after 5 minutes, offering a
link to continue from where it stopped, or as soon as the client
disconnects. Regex matching is aborted at the same deadline to guard against
catastrophic backtracking. The query is carried through all navigation links.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 16:54:19 +09:00
Alex OsborneandClaude Opus 5.5 b1a36159ae Paged log viewer: view a log together with its checkpoint generations
Checkpointing renames each log in place (crawl.log becomes
crawl.log.cp00001-<timestamp>) and starts a new one, so earlier entries
could only be viewed one file at a time. The paged viewer now offers an
"include N checkpoint logs" link (all=y) that shows the log and its
rotated generations in the current launch as one virtual file, oldest
first, with each run of lines labelled with its file.

Positions in the virtual file stay valid across checkpoints, since a
rotated file keeps its place and the new log continues after it, and
they are converted when toggling so the view stays where it was.
Gzipped generations can't be read at arbitrary positions and are
left out.

LogSeries opens the files and fixes their lengths up front, so a
rotation mid-request doesn't change what is read. Paging uses a new
FilteredLineScanner over a LogSeries instead of FileUtils.pagedLines,
which can only read a single file. Results are the same for every
position the navigation links generate; a hand-edited pos in the
middle of a line now starts at that line rather than the next one.
Lines are capped at 256 KiB so a file with few or no newlines can't
exhaust the heap.

Also fixes in the crawl.log rendering: a rotated crawl.log gets the
same status highlighting, status codes unknown to Jetty (e.g. 999) and
numbers too long to be a status no longer fail the page, and the file
name, status titles and navigation links are HTML escaped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 16:46:36 +09:00
Alex Osborne 838523d33a BotBlockDetector: enable in default profiles 2026-10-02 09:24:44 +09:00
Alex Osborne f6a14ac1f6 Implement invoker commands for older browsers like Firefox ESR
Fixes #773
2026-10-01 19:05:53 +09:00
Leslie Bellony b0beb39518 Call tallySourceStats only if the URI is effectively crawled 2026-09-09 17:12:29 +02:00
Leslie Bellony 7b43a91e38 Include Heritrix status code in extended SourceTags report 2026-09-08 11:13:33 +02:00
Alex Osborne b5f32616c8 Fix javadoc warnings 2026-08-25 11:02:15 +09:00
Leslie Bellony 78cc54eb1f Issue #753 verify if bdb is started before generating reports 2026-08-04 18:36:16 +02:00
Leslie Bellony 71bffa42e6 Issue #748 extend SourceTags report with status code 2026-08-03 14:25:53 +02:00
Leslie Bellony 58054b962d Issue #747 add heritrix response codes in ResponseCode report 2026-08-03 11:18:11 +02:00
Alex Osborne 97d8307420 BrowserProcessor: restart the browser automatically if it crashes 2026-07-10 16:12:52 +09:00
Alex Osborne d7676b1b67 BrowserProcessor: log failures continuing intercepted requests 2026-07-10 16:12:47 +09:00
Alex Osborne 1b107f06c1 BrowserProcessor: guard stop() against a partially failed start()
If start() throws after the proxy is started but before the browser
launches, stop() could NPE on the null fields.
2026-07-10 15:59:47 +09:00
Alex Osborne d03a43266a BrowserProcessor: catch tab close failures
If closing the tab in the finally block throws, just log it rather than propagating to the toe thread.
2026-07-10 15:59:46 +09:00
Alex Osborne e276d83f4b BrowserProcessor: avoid NPE when a WebDriverException has no message 2026-07-10 15:48:42 +09:00
Alex Osborne d3da2f9afe BrowserProcessor: handle result.isFailed() in SubresourceRecorder.onComplete() 2026-06-18 23:30:48 +09:00
Alex Osborne 2c9f167ae3 BrowserProcessor: handle recording truncation with length and timeout limits 2026-06-18 23:25:20 +09:00
Alex Osborne 4cb48ec5dc BrowserProcessor: fix content digest and recording limits for subresources 2026-06-18 23:18:23 +09:00
Alex Osborne ff6f98d38d BrowserProcessor: set User-Agent header
Ensures requests made by the browser use the user-agent string configured for the page in the job config.
2026-06-18 23:00:21 +09:00
Alex Osborne e156997dbd MitmProxy: Fix request recording
This fixes a bug where POST requests would get recorded with a doubled header and a bunch of nulls at the end.

It seems that the onRequestContent() listener doesn't pass us the right buffer when proxying. The buffer seems to contain client-to-server data and the buffer's limit isn't set. So instead this records the request content by overriding newProxyToServerRequestContent.
2026-06-17 16:56:24 +09:00
Alex Osborne 709ac00dd2 Add PaginationBehavior: repeatedly clicks next-page and extracts links 2026-06-11 13:34:55 +09:00
Alex Osborne 54b95f5e97 groovy profile: remove stray } 2026-06-10 18:23:12 +09:00
Alex Osborne bd34a7a0a9 Log the exception when creating new jobs fails 2026-06-10 17:43:57 +09:00
Alex Osborne d3defaacd2 Merge pull request #732 from internetarchive/remove-foundation-js
Remove obsolete JavaScript libraries
2026-05-12 16:44:16 +09:00
Alex Osborne 56a1fc2198 Remove Internet Explorer compatibility shims
There's no need for these anymore, IE has been sunset.
2026-05-04 08:45:41 +09:00
Alex Osborne 05f5243207 Remove Foundation JS libraries
Foundation 4 has been obsolete since 2013, and these libraries trigger security alerts. We're also not using them to do much, just the modal dialogs, alert close buttons and collapsing the top-bar on mobile. This replaces the modals with the native HTML dialog element and implements the trivial click event handlers needed for hiding and showing things.
2026-05-04 08:45:40 +09:00
Alex Osborne 71ae9a013b Merge pull request #728 from internetarchive/groovy-config-editor
Add Groovy mode to config editor
2026-05-03 12:06:26 +09:00
Alex Osborne 82181f3d9e Add Groovy mode to config editor
Includes bean and property name autocomplete, similar to the XML mode. Completion of short class names auto-inserts import statements as needed.
2026-04-28 09:08:23 +09:00
Alex Osborne 9e43fde053 Recognise Groovy profiles during job discovery 2026-04-27 17:56:23 +09:00
Alex Osborne b1e13c550b Add profile selector to job creation form
This means:

- Job profiles can be used without having to hunt for them among the regular jobs.
- The Groovy default profile can be selected as an alternative to Spring XML one.
- The default profiles can be customised by creating a job profile with the same name.
2026-04-27 17:40:37 +09:00
Alex Osborne 55c1c20295 Remove httpProxyHost and httpProxyPort from fetcher properties after test execution 2026-04-25 21:28:44 +09:00
Alex Osborne 7490998843 Merge pull request #721 from simons-hub/fix/warn-on-precedence-clipping
Log warning when URI precedence exceeds maximum 127
2026-04-06 16:54:56 +09:00
Simon 47632eb4c4 Log warning when URI precedence exceeds maximum 127
BdbMultipleWorkQueues.calculateInsertKey() silently clips precedence
values to 127 for queue ordering. This can cause unexpected behavior
for custom cost assignment policies that return values above 127.

Log a warning when clipping occurs and fix the misleading comment in
CrawlURI.holderCost that said "should not exceed 255" (the actual
limit is 127, not 255).

Fixes #502
2026-03-31 00:37:36 -07:00
Alex Osborne 96737799ed Enable XXE protection when parsing crawl job XML
Heritrix's Spring XML based job config format by design allows arbitrary code execution, so this doesn't realistically make things much safer, but I also don't see any reason we need external entity resolution enabled. I suppose there's a chance this could mitigate a generic automated attack that uses XXE but doesn't target Heritrix or Spring XML specifically.

Fixes #711
2026-03-02 10:11:38 +09:00
Erik Körner c1bedebef1 Add job deletion to API and UI 2026-01-05 14:19:17 +01:00
Erik Körner 107a249a5d Limit height of Exit Java joblist, set auto-scroll behaviour on overflow 2025-12-18 12:26:08 +01:00
Erik Körner 78b76345d0 Update other endpoints to use new JsonMarshaller, fix some edge cases 2025-12-17 23:25:37 +01:00
Erik Körner cf87b045b9 Add /beans endpoint JSON handling 2025-12-17 22:51:37 +01:00
Erik Körner 18b569736c Add org.restlet.ext.json dependency, add application/json REST API responses 2025-12-17 18:02:36 +01:00
Leslie Bellony 8a7db14a51 Add sizeOnDisk on job status (#699) 2025-12-09 15:59:03 +01:00
Manikya Rathore e4523d63ee fix 690 issue Consolidate and log on deleting URIs from BdB database 2025-11-22 18:42:36 +05:30
Alex Osborne c0e5018eef Config editor: implement autocompletion
This adds context-sensitive autocompletion for bean classes, bean ids, property names, basic Spring XML tags and attributes. The bean completions use a new `/engine/beandoc` endpoint that serves the combination of all the `/META-INF/heritrix-beans.json` files generated at compile-time by the heritrix-docgen annotation processor.
2025-10-30 09:53:16 +09:00
Adam Miller 558727ac37 fix handle more bdb shutdown interrupts 2025-10-03 17:17:46 -07:00
Adam Miller 38e7156ef6 fix: don't restore crawlEndTime when resuming from checkpoint. 2025-09-25 10:50:47 -07:00
Adam Miller fd2aed89fb feat: strip invalid chars from xml rest api output 2025-09-11 17:42:07 -07:00