* ActionDirectory.java
support for .seeds, .recover, .include, .schedule, .force files
* CrawlController.java
change default for pause-at-start to true; ensure enters PAUSED rather than PAUSING
* Frontier.java
(importRecoverFormat) support method for .recover, .include, .schedule, .force additions
(considerIncluded) change to accept CrawlURI
* AbstractFrontier.java
(importRecoverFormat) support method for .recover, .include, .schedule, .force additions
* BdbFrontier.java
remove too-strict assertion
* FrontierJournal.java
rename recover-log 'frontier.recover.gz' (so that '.recover' is a dot-suffix)
* WorkQueueFrontier.java
(considerIncluded) change to accept CrawlURI
* CrawlerJournal.java, IoUtils.java
move getBufferedReader to IoUtils
* FrontierJournal.java
add new 'included' tag for log lines ("Fi")
* (others)
update for relocated method
* CachedBdbMap.java
discard code to open/manage own BDB Environments (and related constructor)
cleanup comments/FIXMEs
* BdbModule.java
return CachedBdbMap, rather than more general ConcurrentMap, in case requester wants it
* CachedBdbMapTest.java
get instance via BdbModule (to manage env/classCatalog)
[HER-1660] allow adding seeds, recovery-journal-like URI data at any point in crawl via 'action directory'
[HER-769] Support case where millions of seeds.
* AbstractFrontier.java
move sourceTagSeeds setting to SeedModule
trigger and schedule seeds via SeedModule's announcements and SeedListener protocols
* StatisticsTracker.java
implement SeedListener, use processedSeedRecords as list of all seeds
* ProcessorURI.java, DefaultProcessorURI.java
(getURI) convenience accessor for plain URI string
* SurtPrefixedDecideRule.java
use seeds as prefixes via SeedListener notifications, rather than separate scan of seeds source text
* SeedListener.java
simplify to notifications of new seed, and nonseed text mixed with seeds (as with SURT directives)
* SeedModule.java
methods for SeedListener interface, announceSeeds, actOn(File)
* TextSeedModule.java
eliminate interators; add methods for announcing seeds and scanning seed text sources (including arbitrary Files)
* PrefixSet.java, SurtPrefixSet.java
derive from ConcurrentSkipListSet; eliminate now-unneeded extra synchronization and clone()s
offer considerAsDirective() for direct application of directive Strings
* CachedBdbMap.java
remove debug output
* TestUtils.java
add info logging
* CachedBdbMapTest.java
make more robust against prior heap usage, different platforms
* CachedBdbMap.java
do expunge on put(), replace()
add low-memory-sensitive 'canary' to force expunge even if otherwise untriggered
* CachedBdbMapTest.java
add test of idle expunge in low-memory conditions
* CrawlURI.java
(getURI) direct access to String URI (don't reuse toString() functionally)
* SeedRecord.java
support for updating record with later/repeat report
* StatisticsTracker.java
remove put()s for processedSeedRecords, hostsLastFinished
* BdbFrontier.java
(getQueueFor) remove put, do in concurrent-compliant manner
* LongToIntConsistentHash.java + Test
consistent-hashing utility class
* URIAuthorityBasedQueueAssignmentPolicy.java
shared superclass for the hostname and surtauthority QAPs
setting for 'deferToPrevious' -- avoid changing assignments
setting for 'parallelQueues' -- when > 1, consistent-hash URIs across that many separate queues (with numerical suffix)
* SurtAuthorityQueueAssignmentPolicy.java, HostnameQueueAssignmentPolicy.java
refactor to derive from URIAuthorityBasedQueueAssignmentPolicy
* QueueAssignmentPolicy.java
add Apache license
* Hop.java
add String representation for convenience
* RecordingInputStream.java
make sure one exception (intermittently seen on Windows re: 'mapped section') doesn't leave things in foul state for next try in same thread
* ConfigPathConfigurer.java
work in two passes: one (BeanPostProcessor) collects all candidates for fixup, two (ApplicationListener) actually performs fixup
suppress fixup of own base path, preventing circulare ref/StackOverflow
* ConfigPath.java
change sense of 'merge' to merge into receiver
* SurtPrefixedDecideRule.java
use ConfigFile, use 'merge' in assignment
* CrawlerLoggerModule.java
use merge in assignment
* CrawlJob.java
(getBeanpathTarget) utility method
* JobResource.java
respect unset ConfigPaths
* ConfigPathEdito.java, ConfigFileEditor.java
better comments, ConfigFile variant
* CrawlController.java
(isFinished) added
* CrawlJob.java
rename reset to teardown
wait for CrawlController to report finished before discarding AppContext
if not finished in time, fail reporting 'false'
* JobResource.java
show Flash if teardown did not succeed; user must retry
* BdbUriUniqFilter.java
don't throw DatabaseException from simple accessor
rename 'discard' to 'teardown'
enable browse/edit of paths outside jobdir via new /anypath/ service
* EngineResource.java, Engine.java
enable addition of job dirs from outside main jobs directory
* EngineApplication.java
register new /engine/anypath service for outside-job-dir paths
Fix for [HER-1568] ARCReader fails on some newer Alexa ARC files
* ARCConstants.java
added ArcRecordErrors enum for known reading errors
* ARCReaderFactory.java
only check for 'LX' extension in GZIP extra field
* ARCRecord.java
added errors, hasErrors(), and getErrors()
* RecordingOutputStream.java
close ReplayInputStream used to create GenericReplayCharSequence
use 'UTF-8' and not possibly-varying JVm default encoding when none specified
* HOWTO-Launch-Heritrix.txt
removed; info mostly about H2 variants
* README.txt
updated wiki link, launch info to match H3
* LICENSE.txt
updated to Apache2