* PersistProcessor.java
add readOnly option for environment only opened for copying-from
* PersistLoadProcessor.java
make start() safe to call redundantly (per Lifecycle contract)
* BdbModule.java
fail-fast if not started properly; stop with fewer errors in problem situations
* CrawlController.java
continue past runtimeException on stop
* CrawlJob.java
less redundant error-reporting
* AbstractFrontier.java
consider PAUSE or FINISH valid targets after unexpected exception
* StatisticsTracker.java
avoid promoting runtimeException in report-generating up
* PersistOnlineProcessor.java
delegate DB-closing to BdbModule to avoid double-close
* BeanBrowseResource.java
coerce edits to type of value they're replacing; makes for safer mid-crawl edits of values in maps
* JobRelatedResource.java
improve rendering/clickability of map, indexed objects
* SurtPrefixedDecideRule.java
allow seeds to be null (as when rule is intentionally completely
independent of seeds; may also need to disqualify rule from
Spring autowiring in such a case)
* SurtPrefixSet.java
refactor to allow add-directives under outside control
* SurtPrefixedDecideRule.java
extend nonseedLine() mechanism to trigger off '+' or '-' depending on decision being ACCEPT or REJECT
* ActionDirectory.java
support for .seeds, .recover, .include, .schedule, .force files
* CrawlController.java
change default for pause-at-start to true; ensure enters PAUSED rather than PAUSING
* Frontier.java
(importRecoverFormat) support method for .recover, .include, .schedule, .force additions
(considerIncluded) change to accept CrawlURI
* AbstractFrontier.java
(importRecoverFormat) support method for .recover, .include, .schedule, .force additions
* BdbFrontier.java
remove too-strict assertion
* FrontierJournal.java
rename recover-log 'frontier.recover.gz' (so that '.recover' is a dot-suffix)
* WorkQueueFrontier.java
(considerIncluded) change to accept CrawlURI
* CrawlerJournal.java, IoUtils.java
move getBufferedReader to IoUtils
* FrontierJournal.java
add new 'included' tag for log lines ("Fi")
* (others)
update for relocated method
* CachedBdbMap.java
discard code to open/manage own BDB Environments (and related constructor)
cleanup comments/FIXMEs
* BdbModule.java
return CachedBdbMap, rather than more general ConcurrentMap, in case requester wants it
* CachedBdbMapTest.java
get instance via BdbModule (to manage env/classCatalog)
[HER-1660] allow adding seeds, recovery-journal-like URI data at any point in crawl via 'action directory'
[HER-769] Support case where millions of seeds.
* AbstractFrontier.java
move sourceTagSeeds setting to SeedModule
trigger and schedule seeds via SeedModule's announcements and SeedListener protocols
* StatisticsTracker.java
implement SeedListener, use processedSeedRecords as list of all seeds
* ProcessorURI.java, DefaultProcessorURI.java
(getURI) convenience accessor for plain URI string
* SurtPrefixedDecideRule.java
use seeds as prefixes via SeedListener notifications, rather than separate scan of seeds source text
* SeedListener.java
simplify to notifications of new seed, and nonseed text mixed with seeds (as with SURT directives)
* SeedModule.java
methods for SeedListener interface, announceSeeds, actOn(File)
* TextSeedModule.java
eliminate interators; add methods for announcing seeds and scanning seed text sources (including arbitrary Files)
* PrefixSet.java, SurtPrefixSet.java
derive from ConcurrentSkipListSet; eliminate now-unneeded extra synchronization and clone()s
offer considerAsDirective() for direct application of directive Strings
* CachedBdbMap.java
remove debug output
* TestUtils.java
add info logging
* CachedBdbMapTest.java
make more robust against prior heap usage, different platforms
* CachedBdbMap.java
do expunge on put(), replace()
add low-memory-sensitive 'canary' to force expunge even if otherwise untriggered
* CachedBdbMapTest.java
add test of idle expunge in low-memory conditions
* CrawlURI.java
(getURI) direct access to String URI (don't reuse toString() functionally)
* SeedRecord.java
support for updating record with later/repeat report
* StatisticsTracker.java
remove put()s for processedSeedRecords, hostsLastFinished
* BdbFrontier.java
(getQueueFor) remove put, do in concurrent-compliant manner
* LongToIntConsistentHash.java + Test
consistent-hashing utility class
* URIAuthorityBasedQueueAssignmentPolicy.java
shared superclass for the hostname and surtauthority QAPs
setting for 'deferToPrevious' -- avoid changing assignments
setting for 'parallelQueues' -- when > 1, consistent-hash URIs across that many separate queues (with numerical suffix)
* SurtAuthorityQueueAssignmentPolicy.java, HostnameQueueAssignmentPolicy.java
refactor to derive from URIAuthorityBasedQueueAssignmentPolicy
* QueueAssignmentPolicy.java
add Apache license
* Hop.java
add String representation for convenience