* profile-crawler-beans.cxml
add comments for all unstated default values an operator might want to change
* (many)
reorder, refactor, rename to better minimize/match simple configuration
* defaults.xml
update bundled profile for new chains refactoring
* CrawlController.java
move to three processor-chains rather than one
* Frontier.java
(loadSeeds) removed
* ToeThread.java
delegate most processing-loop to FetchChain and DispositionChain
* AbstractFrontier.java, WorkQueueFrontier.java
move policies/calculation out to processors
* CandidatesProcessor.java
new processor for DispositionChain that runs every outlink through CandidateChain
* CrawlStateUpdater.java -> DispositionProcessor.java
rename, expand to prep CrawlURI for frontier
* FrontierScheduler.java
deleted; use CandidatesProcessor/CandidateChain
* LinksScoper.java
deprecated; use CandidatesProcessor/CandidateChain/CandidateScoper
(only temporarily retained for ease of comparison)
* CandidateScoper.java
simple single-URI scope-testing for CandidateChain
* FrontierPreparer.java
precalculate all frontier-policies in CandidateChain, before scheduling
* PreconditionEnforcer.java
ProcessorURI->CrawlURI; take on some prerequisite preparation previously deferred to elsewhere
* ProcessorsReport.java
update for 3-chains of Processors
* SheetOverlaysManager.java
(applyOverridesTo) moved here for broader use
* CandidateChain.java, FetchChain.java, DispositionChain.java
role-specific subclasses of ProcessorChain (suitable for type-based autowiring)
* CrawlURI.java
new fields/accessors of use to new chains/frontier
* PostProcessor.java
deleted; skip-to-'postprocessing' is now skip-to-end-of-chain
* ProcessorChain.java
take-on control loop formerly in ToeThread
* ProcessResult.java
absorb ProcessStatus
eliminate problematic STUCK result
* ArchiveUtils.java
(checkContinue) utility method relocated from ToeThread for converting interrupt status to exception
* ClassKeyMatchesRegExpDecideRule.java
properly set up crawlController property for auto-wiring
* (many)
update license notice to Apache where appropriate
eliminate ProcessorURI, DefaultProcessorURI in favor of CrawlURI (now in modules package)
other comment/warning/unused code cleanup
* BdbModule.java
split openDatabase() to openManagedDatabase() (auto-closed) and plain openDatabase() (caller-closes)
set TempStoredSortedMap to use unmanaged openDatabase
* (others)
use openManagedDatabase()
* BdbModule.java
add temp-stored-map service method
remove unused secondary-db support
* TempStoredSortedMap.java
StoredSortedMap that can destroy its underlying database after temporary use
* StatisticsTracker.java
update various sorted-by-decreasing-frequency methods to use temp StoredMaps, with duplicate keys that are negative counts
* (Multiple)Report.java
update to use new duplicate-keyed frequency maps
* PreconditionEnforcer.java, CrawlServer.java
in testing found some ordering issues affecting robots-scheduling if parallelQueues>1
reordered/refactored to put more responsibility in CrawlServer to resolve
* CrawlURI.java
(getPolicyBasisUURI) return either own UURI, or -- for prereqs -- via UURI, as basis for overlays/policy calculations
* URIAuthorityBasedQueueAssignmentPolicy.java
use getPolicyBasisUURI for queue-name decisions
* SheetOverlaysManager.java
use getPolicyBasisUURI for surt-based overlay decisions
alternate simplified implementation of our object-cache need
* ObjectIdentityCache.java
new interface, far less than (Concurrent)Map, for big seems-like-in-memory object-cache
* Supplier.java
trivial interface for deferred-provision of new instance
* CachedBdbMap.java
implement ObjectIdentityCache, mainly via getOrUse()
* ObjectIdentityMemCache.java
trivial ConcurrentHashMap-based all-in-memory ObjectIdentityCache implementation
* ObjectIdentityBdbCache.java
BDB-backed ObjectIdentityCache implementation, carved from CachedBdbMap
* BdbModule.java
refactor utility methods to offer CachedBdbMap or ObjectIdentityBdbCache instances, via source toggle
* EnhancedEnvironment.java
convenience test-environment method
* (others)
update to use ObjectIdentityCache/ObjectIdentityMemCache in place of ConcurrentMap/ConcurrentHashMap
* BdbModule.java
fail-fast if not started properly; stop with fewer errors in problem situations
* CrawlController.java
continue past runtimeException on stop
* CrawlJob.java
less redundant error-reporting
* AbstractFrontier.java
consider PAUSE or FINISH valid targets after unexpected exception
* StatisticsTracker.java
avoid promoting runtimeException in report-generating up
* PersistOnlineProcessor.java
delegate DB-closing to BdbModule to avoid double-close
* BeanBrowseResource.java
coerce edits to type of value they're replacing; makes for safer mid-crawl edits of values in maps
* JobRelatedResource.java
improve rendering/clickability of map, indexed objects
* ActionDirectory.java
support for .seeds, .recover, .include, .schedule, .force files
* CrawlController.java
change default for pause-at-start to true; ensure enters PAUSED rather than PAUSING
* Frontier.java
(importRecoverFormat) support method for .recover, .include, .schedule, .force additions
(considerIncluded) change to accept CrawlURI
* AbstractFrontier.java
(importRecoverFormat) support method for .recover, .include, .schedule, .force additions
* BdbFrontier.java
remove too-strict assertion
* FrontierJournal.java
rename recover-log 'frontier.recover.gz' (so that '.recover' is a dot-suffix)
* WorkQueueFrontier.java
(considerIncluded) change to accept CrawlURI
* CrawlerJournal.java, IoUtils.java
move getBufferedReader to IoUtils
* FrontierJournal.java
add new 'included' tag for log lines ("Fi")
* (others)
update for relocated method
[HER-1660] allow adding seeds, recovery-journal-like URI data at any point in crawl via 'action directory'
[HER-769] Support case where millions of seeds.
* AbstractFrontier.java
move sourceTagSeeds setting to SeedModule
trigger and schedule seeds via SeedModule's announcements and SeedListener protocols
* StatisticsTracker.java
implement SeedListener, use processedSeedRecords as list of all seeds
* ProcessorURI.java, DefaultProcessorURI.java
(getURI) convenience accessor for plain URI string
* SurtPrefixedDecideRule.java
use seeds as prefixes via SeedListener notifications, rather than separate scan of seeds source text
* SeedListener.java
simplify to notifications of new seed, and nonseed text mixed with seeds (as with SURT directives)
* SeedModule.java
methods for SeedListener interface, announceSeeds, actOn(File)
* TextSeedModule.java
eliminate interators; add methods for announcing seeds and scanning seed text sources (including arbitrary Files)
* PrefixSet.java, SurtPrefixSet.java
derive from ConcurrentSkipListSet; eliminate now-unneeded extra synchronization and clone()s
offer considerAsDirective() for direct application of directive Strings