Commit Graph
420 Commits
Author SHA1 Message Date
Kenji Nagahashi 64f24d4c66 HER-1993: In JobResource, write error message instead of throwing RuntimeException when it failed to read
job log or crawl log.
2012-03-09 16:48:11 -08:00
Kenji Nagahashi f05287c8ca restored deleted space after header "Crawl Log" 2012-03-09 16:29:52 -08:00
Kenji Nagahashi adc82c1c6f Second round of HTML cleanup:
- added missing <html> and <body>
- fixed non-working CSS link, added CSS link to all pages
- replaced layout by <br> and &nbsp; with structural tags with CSS
- removed superfluous output, notably in BeanBrowseResource
- <code>-ify variable names in ScriptResource
- added special handling of NaN in doubleToString()
2012-03-03 00:46:56 -08:00
Kenji Nagahashi 46c335c8a3 BeanBrowseResource: fixed malformed HTML. 2012-02-29 13:22:33 -08:00
Noah Levitt 7af7cbf18c Merge branch 'master' of github.com:internetarchive/heritrix3 2012-02-10 15:51:15 -08:00
Noah Levitt db801b3e0f Fix problems handling crawl job paths with characters that require url-encoding, since we sometimes refer to them with file: urls
* PathSharingContext.java, CrawlJob.java
    use java.net.URI to convert paths to and from file: urls
2012-02-10 15:48:09 -08:00
Noah Levitt d83ba3070b Merge branch 'master' of https://github.com/travisfw/heritrix3 2012-02-09 16:43:19 -08:00
Noah Levitt cfac7ead91 Fix for HER-1985 H3: SurtPrefixDecideRule forgets learned/seed-derived/seed-directive SURT prefixes in checkpoint-resume
* SurtPrefixedDecideRule.java
    make class implement Checkpointable, save surt prefixes to json on checkpoint and load them on recover
* profile-crawler-beans.cxml
    make main SurtPrefixedDecideRule a top-level bean so that it can be checkpointed
* Checkpoint.java, CheckpointService.java
    add some FINE level logging during checkpointing and recovery
2012-02-09 10:05:07 -08:00
Travis Wellman e90fc2965b Sort script engines alphabetically so that they appear consistently.
Note that the name of the engine may be different than the name of the language
it parses. A good example is the Rhino script engine that parses javascript.
2012-01-26 00:37:19 -08:00
Noah Levitt a652a9afd0 * profile-crawler-beans.cxml
commented-out default value for new CandidatesProcessor settings processErrorOutlinks
2012-01-20 11:30:24 -08:00
Noah Levitt 4e05c89156 Merge branch 'master' of github.com:internetarchive/heritrix3 2012-01-19 13:41:14 -08:00
Noah Levitt 803f9ca979 HER-1984 save script state - implement by adding a map to the application context for arbitrary use, and make sure the app context is available in all scripting environments
* PathSharingContext.java
    new member variable ConcurrentHashMap data and accessor getData()
* ScriptedProcessor.java, ScriptedDecideRule.java
    make appCtx available to scripts; also remove unused member sharedMap
* ActionDirectory.java
    formatting fix
2012-01-19 13:31:57 -08:00
gojomo 20502f54d5 add processErrorOutlinks setting to CandidatesProcessor
default false is old behavior, skip candidates-handling of outlinks from response codes <200 and >=400
if true, these outlinks will be treated the same as others (get scope-tested and enqueued)
2012-01-18 01:23:26 -08:00
gojomo 4829160c84 [HER-1981] H3: overlay setting of alternate queue-precedence (via
BaseQueuePrecedencePolicy.basePrecedence) doesn't stick
(CrawlURI) don't null overlayNames in processingCleanup
(WorkQueueFrontier) move queue precedence recalc to only in handleQueue,
not all (post-wake) reenqueues
2012-01-05 23:01:54 -08:00
Noah Levitt abda8a01fb HER-1963 Unvisited hosts being written twice into H3-generated hosts-report.txt
* HostsReport.java
    write() - remove obsolete code described by this no longer correct comment - "StatisticsTracker doesn't know of zero-completion hosts; so supplement report with those entries from host cache" - StatisticsTracker does know of zero-completion hosts
2011-11-02 13:01:40 -07:00
Noah Levitt 995133dc40 Fix for HER-1962 NPE from missing sheet
* SheetOverlaysManager.java
    getOverlayMap(String) - return null if sheet missing instead of triggering npe
* KeyedProperties.java
    get(String) - check for null return value from getOverlayMap() and log warning
2011-10-27 10:40:30 -07:00
Noah Levitt 239f616c17 * **/.gitignore
dummy files to make git keep these directories which are expected by unit tests
2011-10-21 22:40:44 -07:00
nlevitt 5735839252 Post 3.1.0
* **/pom.xml
   switch version to "3.1.1-SNAPSHOT"
2011-10-21 19:00:37 +00:00
nlevitt 8c99fc947c Prep for 3.1.0 release
* **/pom.xml
    bump version-id to "3.1.0"
* README.txt
    refer to exactly 3.1.0 release notes
2011-10-21 17:08:04 +00:00
nlevitt e9a96ebc12 * profile-crawler-beans.cxml
commented-out ReschedulingProcessor
2011-10-21 01:20:32 +00:00
nlevitt 9fb7b0f5ae Fix for HER-1958 race condition on frontier inactive queues
* WorkQueueFrontier.java
    activateInactiveQueue(), deactivateQueue() - synchronize around updates of inactiveQueuesByPrecedence and highestPrecedenceWaiting
2011-10-13 01:47:25 +00:00
nlevitt 189a557af8 Tweak of last checkin
* Heritrix.java
    instanceMain() - move logging of java vendor/version ahead of version check, since the information could be particularly helpful when version check doesn't pass
2011-10-12 18:32:14 +00:00
nlevitt 61240573b3 * Heritrix.java
instanceMain() - log java vendor, version at startup
2011-10-12 18:20:27 +00:00
nlevitt 73e6d21fd2 * profile-crawler-beans.cxml
fix small typo
2011-10-12 17:49:34 +00:00
nlevitt 10b3686ba8 Fix race condition which can happen with multiple concurrent "launch"es - can end up with duplicate ToePools like before r7268, probably other problems
* CrawlJob.java
    launch() - synchronize this method - sort of a blunt fix but probably safest and shouldn't affect performance
2011-10-07 23:18:43 +00:00
nlevitt 19e94f07d1 Fix race condition encountered by Travis and Adam where two ToeThread pools are created, one on "launch" and one on "unpause"
* CrawlController.java
    requestCrawlResume() - do not create ToePool (creation here appeared to be inherited cruft), paused crawl must already have toe pool
2011-10-07 01:54:29 +00:00
nlevitt b4e6f43052 Unset script engine variables after scripts complete, allowing memory to be freed up sooner - see also r7252
* ScriptResource.java, ActionDirectory.java, ScriptedDecideRule.java
2011-10-06 18:01:01 +00:00
nlevitt a76f9818f9 Fix HER-1955 Some annoying interaction between new creation of latest link and disk full java bean
* DiskSpaceMonitor.java
    checkAvailableSpace() - really do ignore non-existent paths, as comment and log message claim is done
2011-09-30 20:07:15 +00:00
nlevitt e88d1d9df7 CrawlController FINISHED doesn't guarantee we're ready for shutdown, so add new boolean and event for that purpose.
Elaboration: CrawlJob needs all beans to receive FINISHED event before teardown. Can't easily guarantee CrawlJob receives the event last. Actually, spring provides a way, org.springframework.core.Ordered, but to make a bean receive events last, *every other bean* must implement Ordered. So this way is simpler.
* CrawlController.java
    new member variable isStopComplete and event StopCompleteEvent 
* CrawlJob.java
    rely on cc.isStopComplete() or StopCompleteEvent to indicate ready for teardown
2011-09-29 17:22:34 +00:00
nlevitt de9975ee23 * CrawlJob.java
instantiateContainer() - do not call teardown() on exception from ac.refresh() - can throw IllegalStateException
    teardown() - put all uses of variable cc inside not-null check
2011-09-29 16:09:44 +00:00
nlevitt 747fbc0a0f Fix for HER-1954 bdb closed at crawl finish instead of teardown -- refactor crawl finish and teardown
* PathSharingContext.java
    start(), doStart(), stop(), doStop() - remove these overloaded methods, because the spring bug they were working around seems to be fixed, not seeing any problems with cyclical dependencies - might be this one https://jira.springsource.org/browse/SPR-7266
    doClose() - remove because superclass version seems to work fine (this method generally wasn't being called anyway, though it would be now with other changes in this checkin)
* CrawlJob.java
    refactor teardown to call close() on the ApplicationContext, which calls destroy() on any beans that implement DisposableBean - this is now the way to have beans do stuff at teardown
* CrawlController.java
    send FINISHED crawl state event after calling appCtx.stop() so that isFinished() can indicate ready-ness for teardown
* BdbModule.java
    move close() to teardown, i.e. implementation of DisposableBean.destroy(); remove shutdown hook and rely on teardown; related tweaks
* WorkQueueFrontier.java, CrawlerLoggerModule.java, BdbUriUniqFilter.java
    move close() to teardown
* CrawlMapper.java, AbstractFrontier.java, PreloadedUriPrecedencePolicy.java, FetchWhois.java, FetchHTTP.java, PersistLogProcessor.java, WriterPoolProcessor.java
    add comments about cleanup that maybe should wait until teardown
* UriUniqFilter.java
    remove incorrect(?) comment
2011-09-29 00:39:55 +00:00
nlevitt de856f10a6 Fix possible memory leak (encountered in h1, see r7255)
* ToeThread.java
    run() - set local variable curi=null when finished with it, because the jvm seems to be holding on to it even after the containing block is finished
2011-09-28 00:22:53 +00:00
nlevitt f704ca5bc7 Fix HER-1953 all api xml elements are camelcase except sizeTotalsReport fields
* CrawledBytesHistotable.java
    camelcase notModified, dupByHash, notModifiedCount, dupByHashCount, novelCount
* CrawlJob.java
    sizeTotalsReportData() - camelcase totalCount
2011-09-28 00:19:44 +00:00
nlevitt 6adc3c285f * Heritrix.java
usage()
        -        out.print("Your arguments were: "+StringUtils.join(args, ' '));
        +        out.println("Your arguments were: "+StringUtils.join(args, ' '));
2011-09-27 04:08:58 +00:00
nlevitt f242522efc Fix for HER-1950 URIAuthorityBasedQueueAssignmentPolicy KeyedProperties shadows superclass QueueAssignmentPolicy KeyedProperties, breaking sheet overlay of properties defined in superclass
* URIAuthorityBasedQueueAssignmentPolicy.java
    remove field kp, accessor getKeyedProperties(); instead inherit these from superclass QueueAssignmentPolicy
2011-09-20 23:52:59 +00:00
nlevitt 54aca17757 Removing very old abandoned line of development, package org.archive.extractor 2011-09-14 21:34:26 +00:00
nlevitt 4eef2d9902 Revert change in r7240 - different return type makes recovery from checkpoints from before the change fail
* CrawlController.java
    getState() - return Object
2011-09-13 00:48:06 +00:00
nlevitt 5a4e72d902 Couple of synchronizations for HER-1943 WorkQueue: inconsistent synchronization using fields active, lastCost, peekItem, wakeTime
* WorkQueue.java
    considerActive(), unpeek() - make synchronized
2011-09-12 22:49:32 +00:00
nlevitt 48b0ddfb03 Fix for HER-1942 CheckpointService: inconsistent synchronization around use of isRunning. Doubtful if isRunning thing is a real problem, but it looks like a good idea to synchronize this method anyway, in case there are concurrent calls to it.
* CheckpointService.java
    setRecoveryCheckpointByName() - mark synchronized
2011-09-12 22:00:54 +00:00
nlevitt f20b3eee30 Fix HER-1935 Many calls to File.mkdirs() and other file/dir methods don't check the return value. Also fix bug where pointless empty directories were created in scratch dir.
* BasicProfileTest.java, SelfTestBase.java, CrawlControllerTest.java, PrecedenceLoader.java, MigrateH1to3Tool.java, CrawlerLoggerModule.java, StatisticsTracker.java, CheckpointUtils.java, BdbUriUniqFilter.java, ARCWriterProcessorTest.java, WARCWriterProcessorTest.java, PersistProcessor.java, WriterPoolProcessor.java, PrefixFinderTest.java, StoredQueueTest.java, FileUtilsTest.java, ObjectIdentityBdbManualCacheTest.java, ObjectIdentityBdbCacheTest.java, ObjectPlusFilesOutputStream.java, TestUtils.java, TmpDirTestCase.java, Engine.java
    Replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir). In these cases the calls were either already in a spot where the possible IOException would be handled appropriately, or the line was trivially moved into such a block.
* ActionDirectory.java, Engine.java
    replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
* BdbModule.java
    setup() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException" - conveniently the place where this method is called was already in a try block that catches IOException
* Checkpoint.java
    generateFrom() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException"
* CheckpointService.java
    move call to Checkpoint.generateFrom() inside existing try block since it now can throw IOException
* Recorder.java
    ensure(File) - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
    new Recorder(File,String,int,int) - call ensure() on the correct object, the containing directory; and remove redundant call to ensure()
2011-09-12 20:12:10 +00:00
nlevitt 221e63dbd0 Fix for HER-1921 BucketQueueAssignmentPolicy: absolute value of hashCode can be Integer.MIN_VALUE.
* BucketQueueAssignmentPolicy.java
    getClassKey() - cast value to long before taking absolute value
2011-09-12 17:02:54 +00:00
nlevitt 24f4b715ec Various little cleanups including several discovered by Aaron using an automated tool.
* WorkQueueFrontier.java
   (HER-1925) avoid 2 possible null pointer dereferences
* CrawlController.java
    getState() - change declared return type from Object to State
* CrawlerLoggerModule.java
    (HER-1932) remove unused, unset field "reports"
* BdbCookieStorage.java
    elide pointless extra variable
* FetchHTTP.java, DownloadURLConnection.java, ProcessUtils.java
   (HER-1933) use Arrays.toString() for logged arrays
* CrawlServer.java
    - remove unused field robotstxtChecksum
    - updateRobots() - avoid reinventing existing utility class InstanceofPredicate
* ExternalGeoLookupInterface.java
    (HER-1938) extend Serializable, since ExternalGeoLocationDecideRule is declared Serializable and has a ExternalGeoLookupInterface field
* DecideRuleSequence.java
    (HER-1937) make field fileLogger transient
* PersistLogProcessor.java
    (HER-1924) remove field recoveryCheckpoint shadowing same field in superclass Processor
* ExtractorUniversal.java
    (HER-1920) use return value of potentialTLD.toLowerCase() as it appears was intended
* ARCWriterProcessor.java
    (HER-1928) avoid possible null pointer dereference
* S3URLConnection.java
    (HER-1933) set S3ServiceException as cause of rethrown IOException, and do not put stacktrace array in message
* ObjectIdentityMemCache.java
    (HER-1929) avoid possible null pointer dereference
* .classpath
    more source jar references, other cleanup
2011-09-12 02:12:44 +00:00
nlevitt e38f0a73ba * profile-crawler-beans.cxml
<bean id="fetchHttp" class="org.archive.modules.fetcher.FetchHTTP">
    +  <!-- <property name="useHTTP11" value="false" /> -->
2011-09-11 17:02:32 +00:00
nlevitt d3a6dc1f58 Fix for HER-1946 Mid-Crawl Adjustment of frontier.queueTotalBudget does not propagate to retired queues.
* WorkQueueFrontier.java
    findEligibleURI() - propagate queueTotalBudget (also sessionBudget) to ensure the settings are current right before checking if queue is over budget
2011-09-09 23:23:15 +00:00
nlevitt c70f7ccbc0 * profile-crawler-beans.cxml
remove reference to defunct frontier setting "holdQueues", pointed out by Kristinn Sigurðsson
2011-08-08 23:42:47 +00:00
nlevitt 28aa2ffede * **/pom.xml
improve formatting (no substantive changes at all -- try "svn diff -x -ub -c7232")
2011-08-05 17:46:57 +00:00
nlevitt 178fdd92b6 * **/pom.xml
remove references to unused, defunct agilejava maven repository (had been used for xdoc stuff)
2011-08-05 17:37:16 +00:00
nlevitt ada8cbd198 Fix bug reported by Travis where in certain corner cases where urls are processed in a particular order, a seed can be incorrectly tagged with S_ROBOTS_PREREQUISITE_FAILURE
* DispositionProcessor.java
    innerProcess() - do not update server's robots info if robots fetch has been deferred
2011-08-04 01:32:09 +00:00
nlevitt 706869e586 Fix handling of server error on robots.txt - change to what appears to be the intended behavior
* DispositionProcessor.java
    innerProcess() - call server.updateRobots(curi) after special ignore-robots short-circuiting of retry cycle rather than before, so that updateRobots() deemed-not-found handling applies
* FetchHTTP.java
    call curi.setHttpMethod() right after creating the HttpMethod, instead of when processing the response; this is needed so that updateRobots() check of fetchType works properly (before this change, fetchType was "UNKNOWN" on server error)
* PreconditionEnforcer.java, CrawlServer.java
    fix typos
2011-08-02 23:28:20 +00:00
nlevitt 3cfc315de9 Fix bug reported by Travis where each url in action directory ".include" file is reported as "not canonicalized" in heritrix_out, and some end up crawled in spite of being considered included
* WorkQueueFrontier.java
    considerIncluded(CrawlURI) - call FrontierPreparer.prepare(CrawlURI) so that CrawlURI is canonicalized etc before being considered included (similar to what schedule() does)
* CrawlURI.java
    getCanonicalString(), getPolitenessDelay() - use logger instead of System.err
2011-08-02 01:07:11 +00:00