Commit Graph
827 Commits
Author SHA1 Message Date
Noah Levitt 4524f24dc7 Merge branch 'warc-writer-chain' into ait-qa
* warc-writer-chain:
  extract watch page links from youtube playlists
  fix non-playlist case (oops!)
  be consistent and null-safe with concurrentTo
2019-11-15 15:54:31 -08:00
Noah Levitt 1b95453748 be consistent and null-safe with concurrentTo 2019-10-15 10:50:26 -07:00
Noah Levitt 5a813197c2 Merge branch 'warc-writer-chain' into ait-qa
* warc-writer-chain:
  write youtube-dl json to the warc
  rename method so as not to conflict with Processor
  make WARCRecordBuilder an interface
  revisits only for http and ftp
  accept any BaseWARCWriterProcessor
  use WARCWriterChainProcessor for these tests
  default chain in code
  oops, handle https too
  same test as for WARCWriterProcessor
  configurable warc writer chain
2019-06-12 17:42:20 -07:00
Noah Levitt bff33e00e6 write youtube-dl json to the warc
ExtractorYoutubeDL implements WARCRecordBuilder
2019-06-12 17:40:57 -07:00
Noah Levitt 5177b2b6da rename method so as not to conflict with Processor 2019-06-12 15:17:33 -07:00
Noah Levitt 9ddd281ef2 make WARCRecordBuilder an interface
this way other classes that extend other classes can also implement
WARCRecordBuilder
2019-06-12 15:04:39 -07:00
Noah Levitt 9e67a8dab4 revisits only for http and ftp
fixes NPE trying to write a dns revisit record
2019-06-12 14:58:11 -07:00
Noah Levitt 7c31b0752a use WARCWriterChainProcessor for these tests 2019-06-12 12:56:49 -07:00
Noah Levitt 8c4c443a88 default chain in code 2019-06-12 12:54:00 -07:00
Noah Levitt ece435874e oops, handle https too 2019-06-12 12:53:27 -07:00
Noah Levitt 870d84740a same test as for WARCWriterProcessor
doesn't test that much though
2019-06-12 11:04:02 -07:00
Noah Levitt 9435f761a6 configurable warc writer chain
exercised only lightly at this point
2019-06-11 13:45:51 -07:00
Noah Levitt aec1443e56 Merge branch 'ydl' into ait-qa
* ydl:
  whoops, better spawn the thread first
  read stderr and stdout in separate threads...
  do not drop any `CrawlURI.data` between processing
  nothing private ever
  quiet org.mortbay.log (jetty?) logging
  making everything work
  ExtractorYoutubeDL
  Update changelog for 3.4.0-20190418
  [maven-release-plugin] prepare for next development iteration
  [maven-release-plugin] prepare release 3.4.0-20190418
  set of frontier management changes to support CrawlHQ module
  Remove suffix from warcWriter since it is no longer used.
  Add CHANGELOG; address #233.
2019-05-02 18:27:36 -07:00
Noah Levitt 37fb6f6b7b do not drop any CrawlURI.data between processing
Without this change (or other measures), we sometimes get nulls in the
ExtractorYoutubeDL log for containing page information. We'll run this
on QA for a while and see if it causes any problems.

nlevitt [1:59 PM]
https://github.com/internetarchive/heritrix3/blob/master/modules/src/main/java/org/archive/modules/CrawlURI.java#L878
drops some stuff from `CrawlURI.data` after processing a uri, even if it needs to be processed again
there is a list of keys that shouldn’t be dropped (`persistentKeys`), but it is final and private
so if you’re writing your own heritrix module and you want to keep some information in CrawlURI.data, it usually works, except when the url is processed more than once (like when it needs a prereq like robots.txt the first time)
in practice it seems that most data is persisted, that is, most commonly used keys are in `persistentKeys`
in a crawl with pretty standard configuration i’m mostly seeing `prerequisite-uri` dropped and occasionally `fetch-completed-time` and `fetch-began-time` being dropped
i’m highly skeptical of the value of dropping keys at all and i’m tempted to get rid of this entirely, make all the keys persistent in other words
soliciting feedback (edited)

anjackson [2:39 PM]
My immediate reaction is HARD AGREE. It looks like Really Old Code though (https://github.com/internetarchive/heritrix3/blame/7d3eff5269142c77fa4b988396153f4c29d16caa/modules/src/main/java/org/archive/modules/CrawlURI.java#L878)
so the reasons for doing so may have been lost in time.
Hm, looking at usage: https://github.com/internetarchive/heritrix3/blob/a60b2ef3875ad47f57b0c6b3c0b19f86c40a12f7/engine/src/main/java/org/archive/crawler/frontier/WorkQueueFrontier.java#L954-L955
engine/src/main/java/org/archive/crawler/frontier/WorkQueueFrontier.java:954-955

                curi.processingCleanup(); // lose state that shouldn't burden
                                          // retry

I guess there's a concern that there may be state in there that is set during a fetch and may cause problems if the same CrawlURI is deferred?
But I'm not aware of anything in the fetch chain that behaves like that.

nlevitt [3:02 PM]
oh, i missed `CrawlURI.addDataPersistentMember(String)` et al. still...
2019-04-30 16:20:53 -07:00
Noah Levitt 1bd8b713c6 quiet org.mortbay.log (jetty?) logging
There is already a clause for this in logging.properties, but it's using
log4j. It was dumping stack traces every time the client was dubious of
heritrix's self-signed certificate.

Why do we have so many identical log4j.xml's? 🤷‍♂️
2019-04-30 16:17:35 -07:00
Andrew Jackson b3961a2f96 [maven-release-plugin] prepare for next development iteration 2019-04-18 15:36:28 +01:00
Andrew Jackson c7c6141ee1 [maven-release-plugin] prepare release 3.4.0-20190418 2019-04-18 15:36:20 +01:00
Noah Levitt 0aac882eb4 Merge branch 'master' into ait-qa
* master:
  use constant from rethinkdb lib for default port
  Update README.md
  Handle missing closing paren in srcset descriptor
  Teach jericho extractor srcset
  Don't run srcset test against jericho, it doesn't handle it
  Handle commas more compliantly when parsing srcset
  Ensure we start parsing full lines, for #239.
2019-04-01 16:16:03 -07:00
Alex Osborne dd37598470 Revert "Upgrade httpclient to 4.5.7 and handle cookies more compliantly" 2019-03-28 10:14:42 +09:00
Andrew Jackson 03da8c06cf Merge branch 'master' into upgrade-httpclient 2019-03-21 00:06:10 +00:00
Andrew Jackson 044f068d1b Removing outdated test. 2019-03-21 00:06:05 +00:00
Andrew Jackson 629a7adcb6 Disable questionalbe test. 2019-03-20 22:02:58 +00:00
Andrew Jackson 2fbe603878 Avoid deprecated flag. 2019-03-20 22:02:35 +00:00
Andrew Jackson e910bb6a5b Supply an iterator, for #245 2019-03-20 21:18:19 +00:00
Alex Osborne 7d91a1d4ee Handle missing closing paren in srcset descriptor 2019-03-16 16:02:48 +09:00
Alex Osborne 2a34fcffd6 Teach jericho extractor srcset 2019-03-16 15:47:50 +09:00
Alex Osborne e1d93e3308 Don't run srcset test against jericho, it doesn't handle it 2019-03-16 12:35:57 +09:00
Alex Osborne 90c52c3a25 Handle commas more compliantly when parsing srcset
Commas are allowed if they're in the middle of the URL. Consequently:

    srcset="a,b,,c,"   => ["a,b,,c"]
    srcset="a, b,, c," => ["a", "b", "c"]

They occur particularly commonly in data: URLs before the base64 value.

Commas are also allowed in descriptors if they are enclosed by parens:

    srcset="a (b,c),d" => ["a", "d"]

Spec: https://html.spec.whatwg.org/multipage/images.html#parsing-a-srcset-attribute
2019-03-16 11:53:37 +09:00
Noah Levitt b12e34d340 Merge branch 'trough-dedup' into ait-qa
* trough-dedup:
  promote dirty segments at crawl finish
  trough dedup!
  Allow failed lookups to expire, for #234.
  [maven-release-plugin] prepare for next development iteration
  [maven-release-plugin] prepare release 3.4.0-20190207
  As @nlevitt suggestion, a further check.
  [maven-release-plugin] prepare for next development iteration
  [maven-release-plugin] prepare release 3.4.0-20190205
  [maven-release-plugin] prepare for next development iteration
  [maven-release-plugin] prepare release 3.4.0-20190205-2
  [maven-release-plugin] prepare for next development iteration
  Extend to set additional property.
  Use argument syntax.
  Skip tests during release process (covered by CI).
  Set consistent tag, and include contrib.
  Wrong repo spec.
  Add build profile for deployment to Maven Central.
  Avoid headings being treated as lists
  Clarification about APIs.
  Swapped original and link to make maintenance simpler.
  Tidy up markup and links.
  Add synchronized statements for internetarchive/heritrix3#221.
  Add checks to guard against server sending 304 in error, for #229.
  do not checkpoint if crawl job has not started
2019-03-14 15:17:39 -07:00
Andrew Jackson ad9e74d8ad [maven-release-plugin] prepare for next development iteration 2019-02-07 13:52:56 +00:00
Andrew Jackson 83c7044b22 [maven-release-plugin] prepare release 3.4.0-20190207 2019-02-07 13:52:49 +00:00
Andrew Jackson 667cf3ac5e As @nlevitt suggestion, a further check. 2019-02-06 21:03:49 +00:00
Andrew Jackson e8e37751bb Merge branch 'master' into more-robust-fetch-history-processor 2019-02-06 20:47:06 +00:00
Andrew Jackson 317b19e7e2 [maven-release-plugin] prepare for next development iteration 2019-02-05 12:33:08 +00:00
Andrew Jackson 9c7d0c299d [maven-release-plugin] prepare release 3.4.0-20190205 2019-02-05 12:32:57 +00:00
Andrew Jackson ab34a54566 [maven-release-plugin] prepare for next development iteration 2019-02-05 12:29:33 +00:00
Andrew Jackson d1fc5cedb4 [maven-release-plugin] prepare release 3.4.0-20190205-2 2019-02-05 12:29:26 +00:00
Andrew Jackson 096ba4e9c3 [maven-release-plugin] prepare for next development iteration 2019-02-05 12:15:26 +00:00
Andrew Jackson 949a350cd8 Add checks to guard against server sending 304 in error, for #229. 2019-02-04 10:26:12 +00:00
Noah Levitt a7a54ddf2e Merge branch 'deciding-optimization' into ait-qa
* deciding-optimization:
  implement PredicatedDecideRule.onlyDecision()
2018-11-15 11:21:01 -08:00
Noah Levitt fea6241fac implement PredicatedDecideRule.onlyDecision()
DecideRuleSequence already has the optimization that I was looking for,
namely, don't bother evaluating a DecideRule if we know it won't change
the current result. For some reason PredicatedDecideRule, which is a
parent class to most of the decide rules in heritrix, didn't implement
onlyDecision(). Certain crawl configurations could see significant
performance improvement with this change.
2018-11-15 11:13:27 -08:00
Noah Levitt 6bee1cbb6f Merge branch 'hbase-refactor' into ait-qa
* hbase-refactor:
  use non-deprecated hbase api
  Correct spelling mistakes.
  Note feature only applies to forthcoming 3.3 release
  Update API with note about checkpoint launching.
  Default to starting anew if no checkpoints are found.
  Provide a simple hook to restart from the latest checkpoint.
  Move sync to after Null check.
  Attempt to cache dependencies.
  Simplify contrib build.
  Attempt to prevent build time-outs.
  Add synchronisation around host and server stats.
  Avoid overdoing the RAM allocation.
  Ensure all WorkQueue modification as serialised across threads.
  Tests now need more RAM.
  fix exception starting DecideRuleSequence logging
  Fix up tests to account for new Base URI behaviour.
  Only set the BaseURI if not set already.
  Same fix for JerichoExtractorHTML.
  Test case and fix for internetarchive/heritrix3#208.
  Fix link to User Guide
  Fix typos
  Replace API response section with statuscode tag
  Reformat the rest of the API calls
  Add actions to API URLs to distinguish them
  Add explanatory note about documentation
  Try nicer formatting for the first API example
  Format HTTP request lines in API guide
  Move conventions and REST section to the end of the document
  Convert html tables to rst tables
  Add skeleton sphinx docs for readthedocs
  Point at GitHub wiki for latest release info
  Add parameter to allow even distribution for parallel queues.
2018-11-09 15:41:05 -08:00
Noah Levitt a831676196 Merge pull request #209 from ukwa/relative-base-href
HtmlExtractor: allow relative hrefs in the base element
2018-09-11 15:18:48 -07:00
Noah Levitt 9c9d11d272 fix exception starting DecideRuleSequence logging
I don't know why we've never seen this before, and now we suddenly have
a case of it, but this is the exception:

2018-07-23 17:47:01.123 SEVERE thread-2875088 org.archive.crawler.framework.CrawlJob.beansException() Failed to start bean 'scope'; nested exception is java.lang.IllegalStateException: java.nio.file.NoSuchFileException: /1/ait-h3-jobs/8144-20180723162745141/20180723174701/logs/scope.log.lck
org.springframework.context.ApplicationContextException: Failed to start bean 'scope'; nested exception is java.lang.IllegalStateException: java.nio.file.NoSuchFileException: /1/ait-h3-jobs/8144-20180723162745141/20180723174701/logs/scope.log.lck
        at org.springframework.context.support.DefaultLifecycleProcessor.doStart(DefaultLifecycleProcessor.java:169)
        at org.springframework.context.support.DefaultLifecycleProcessor.access$1(DefaultLifecycleProcessor.java:154)
        at org.springframework.context.support.DefaultLifecycleProcessor$LifecycleGroup.start(DefaultLifecycleProcessor.java:335)
        at org.springframework.context.support.DefaultLifecycleProcessor.startBeans(DefaultLifecycleProcessor.java:143)
        at org.springframework.context.support.DefaultLifecycleProcessor.start(DefaultLifecycleProcessor.java:89)
        at org.springframework.context.support.AbstractApplicationContext.start(AbstractApplicationContext.java:1236)
        at org.archive.spring.PathSharingContext.start(PathSharingContext.java:115)
        at org.archive.crawler.framework.CrawlJob.startContext(CrawlJob.java:455)
        at org.archive.crawler.framework.CrawlJob$1.run(CrawlJob.java:428)
Caused by: java.lang.IllegalStateException: java.nio.file.NoSuchFileException: /1/ait-h3-jobs/8144-20180723162745141/20180723174701/logs/scope.log.lck
        at org.archive.crawler.reporting.CrawlerLoggerModule.setupSimpleLog(CrawlerLoggerModule.java:298)
        at org.archive.modules.deciderules.DecideRuleSequence.start(DecideRuleSequence.java:174)
        at org.springframework.context.support.DefaultLifecycleProcessor.doStart(DefaultLifecycleProcessor.java:166)
        ... 8 more
Caused by: java.nio.file.NoSuchFileException: /1/ait-h3-jobs/8144-20180723162745141/20180723174701/logs/scope.log.lck
        at sun.nio.fs.UnixException.translateToIOException(UnixException.java:86)
        at sun.nio.fs.UnixException.rethrowAsIOException(UnixException.java:102)
        at sun.nio.fs.UnixException.rethrowAsIOException(UnixException.java:107)
        at sun.nio.fs.UnixFileSystemProvider.newFileChannel(UnixFileSystemProvider.java:177)
        at java.nio.channels.FileChannel.open(FileChannel.java:287)
        at java.nio.channels.FileChannel.open(FileChannel.java:335)
        at java.util.logging.FileHandler.openFiles(FileHandler.java:478)
        at java.util.logging.FileHandler.<init>(FileHandler.java:344)
        at org.archive.io.GenerationFileHandler.<init>(GenerationFileHandler.java:63)
        at org.archive.io.GenerationFileHandler.makeNew(GenerationFileHandler.java:158)
        at org.archive.crawler.reporting.CrawlerLoggerModule.setupLogFile(CrawlerLoggerModule.java:275)
        at org.archive.crawler.reporting.CrawlerLoggerModule.setupSimpleLog(CrawlerLoggerModule.java:296)
        ... 10 more
2018-07-27 09:49:17 +09:00
Andrew Jackson 8384a59b47 Fix up tests to account for new Base URI behaviour. 2018-07-06 15:15:53 +01:00
Andrew Jackson 8bf23081fa Only set the BaseURI if not set already. 2018-07-05 13:34:17 +01:00
Andrew Jackson 14148cc7a8 Same fix for JerichoExtractorHTML. 2018-07-05 09:40:21 +01:00
Andrew Jackson 42181b8585 Test case and fix for internetarchive/heritrix3#208. 2018-07-04 16:18:46 +01:00
Noah Levitt 74c48659b8 Merge branch 'ari-5630' into ait-qa
* ari-5630:
  catch exceptions scoping outlinks to stop them from derailing processing of the parent url
  fix for test failures in a workspace on NFS-mounted filesystem
  max size for extracted form elements
2018-01-16 17:13:04 -08:00
Kenji Nagahashi a7b7c6cf5a fix for test failures in a workspace on NFS-mounted filesystem
ContentDigestHistoryTest does not close BdbModule. It results in failure to delete bdb directory in following tests.
Added tearDown() method that closes BdbModule.
2017-12-08 17:15:03 -08:00