use heritrix.home, if available, to find default logging.properties (to let those who launch Heritrix from elsewhere save themselves from HttpClient's copious logging)
* StatisticsTracker
move small stat maps to in-memory ConcurrentMaps
update processedSeedRecords for new ObjectIdentityCache shape
move hostsDistribution/hostsBytes/hostsLastFinished tracking to CrawlHosts/hostsCache
* CrawlSummaryReport
gets hosts count from hostsCache
* SeedRecord
implement IdentityCacheable
* FetchStats
remember last-success times
* IdentityCacheable
new interface required of objects stored in ObjectIdentityCaches
* IdentityCacheableWrapper
wrapper for storing arbitrary objects in ObjectIdentityCaches
* ObjectIdentityCache
keys now always Strings
values now always IdentityCacheables
new dirtyKey() method to ensure a key is persisted
* ObjectIdentityBdbCache, ObjectIdentityMemCache
update for new ObjectIdentityCache shape; still using legacy GC magic
* ObjectIdentityBdbManualCache
alternate implementation relying on dirtying to ensure persistence
use Guava library MapMaker for soft memMap; capped dirtyMap
* BdbModule
use ObjectIdentityBdbManualCache by default
up expectedConcurrency default to 64
* BdbModuleTest, ObjectIdentityBdbCacheTest, ObjectIdentityBdbManualCacheTest
update, add tests
* Frontier
FrontierGroup as IdentityCacheable
* AbstractFrontier, WorkQueueFrontier, BdbFrontier
touchup queue/group accessors
make queue instances dirty whenever mutated
* WorkQueue, CrawlHost, CrawlServer
IdentityCacheable support; appropriate makeDirty()s
* ServerCache, DefaultServerCache
update for new ObjectIdentityCache shape
* CrawlURI
MAX_HOPS_DISPLAYED constant, 50
build new CrawlURI pathFromSeed value via extendHopsPath, which shows only MAX_HOPS_DISPLAYED hops, and precedes string with int value of hops unshown
* CrawlURITest
test for extendHopsPath
processing is finished.
* Recorder.java
getReplayCharSequence() - remember ReplayCharSequence, obtain new one if
none remembered or remembered has been closed
endReplays() - close replayCharSequence
* ToeThread.java
run() - call Recorder.endReplays() when finished with uri
* ReplayCharSequence.java
isOpen() - return false if close() has been called
* GenericReplayCharSequence.java, InMemoryReplayCharSequence.java
implement isOpen()
* KeyWordProcessor.java, ExtractorHTML.java, HTTPContentDigest.java,
ExtractorCSS.java, ExtractorXML.java
do not close ReplayCharSequence
* WorkQueue.java
add noteExhausted() for emptied queues
* WorkQueueFrontier.java
add synchronization around checkFutures() activity
use noteExhausted when queue is totally empty
* DecideRuleSequence.java
- use Spring Lifecycle start() to initialize logToFile
- inject logger module as SimpleFileLoggerProvider, since we're in
heritrix-modules and don't have access to CrawlerLoggerModule from
heritrix-engine
* CrawlerLoggerModule.java
- implement SimpleFileLoggerProvider
- setupSimpleLog() - use 'T' instead of '+' between date and time in
timestamp
* SimpleFileLoggerProvider.java
new interface with one method, setupSimpleLog()
* Scoper.java
do not call scope.start()/scope.stop(), these are handled using Spring
Lifecycle now
* DecideRule.java
remove unused, unneeded start()/stop()
* KryoBinding.java
always default to registrationOptional, for now
* StatisticsTracker.java
explicitly set valueClass of sourceDistribution to ConcurrentHashMap
* StoredQueue.java
answer 0 for size() when backing DB closed
* CheckpointService
Made checkpointIntervalMinutes setter accept a long instead of an int (was missed when the class variable was made a long). While it was non-harmful the way it was, this is more in line with Java conventions.
* AbstractFrontier
remove inbound/outbound queues; simplify managerThread
perform findEligible/schedule/receive/finish immediately
* WorkQueueFrontier
eliminate holdQueues setting
synchronize sendToQueue on target queue
make findEligibleUri reentrant; rely on [Blocking|Stored]Queue thread-safety
make arriving at front of ready trigger for 'active'/budget-session
simplify inactiveQueues management
simplify waking from overflow
* WorkQueue
redefine 'active' as in-budget-session
rename 'held' as 'managed'
synchronize major methods, relying on allQueues cache that operations on same intended queue use identical instance
split over-budget test to isOverSessionBudget and isOverTotalBudget
have toString show classKey
* BdbFrontier
checkpoint fixes: remember nextOrdinal, rotate frontier-recover log, reset queues on recovery
consistencyCheck method for probing state during stress tests
* RecordingInputStream.java
check for interrupt on each socket-timeout
* CrawlController.java
on second requestCrawlStop, interrupt threads via ToePool.cleanup
* ToePool.java
adjust for new earlier/repeated cleanup
* ToeThread.java
better warnings/recovery on forced-interrupts
* UriProcessingFormatter.java
allow pre-caching of a specific LogRecord's formatted version, outside the synchronized publish()
* GenerationFileHandler.java
force a format before publish(), to benefit above (a small hit in other cases where it's redundant)
* (many)
distinction between these two constant-collecting classes was fuzzy (to the point that they already referred to each other), and they were both in same subproject/package already as well; so, merged constants into the larger, older class
* WorkQueueFrontier
(findEligibleURI) avoid outbound.capacity-sensitive activation, which under a race created by recent changes led to infinite recursion here
* AbstractFrontier
(next) when nothing is immediately ready, try adding one-at-a-time to outbound, rather than fillOutbound()
* AbstractFrontier
(drainInbound) don't assume current thread will be able to take() full batch; poll() and exit early if other threads have drained queue first
* AbstractFrontier, WorkQueueFrontier
synchronize around all InEvent/findEligibleURI actions, so that drainInbound and fillOutbound may be called outside managerThread
whenever a ToeThread would block on enqueue() or next(), try the appropriate catch-up method before blocking
remove assertions no longer true with activity happening outside managerThread
* AutoKryo
extension of Kryo to allow classes to control their own registration, trigger registration of associated classes, and deserialize classes without no-arg constructors
* KryoBinding
binding for use with BDB that uses AutoKryo serialization for a 2X-4X reduction in byte[] size
* UURI
improved serialization via Externalizable and Kryo's CustomSerialization methods
* BdbModule
discard deprecated CachedBDBMap option
(getObjectCache) extend with both declaredClass and valueClass (for when map values are specializations of the declared type, as with frontier.allQueues)
adjust type declarations
* ObjectIdentityBdbCache
use KryoBinding rather than SerialBinding
* CachedBdbMapTest
discarded
* BdbFrontier, BdbServerCache, StatisticsTracker
adjust type declarations, objectCache creation
* BdbMultipleWorkQueues
use KryoBinding rather than (Recycling)SerialBinding
* BdbWorkQueue, CrawlServer, CrawlHost, CrawlURI
add autoregister support
* LinkContext
public for kryo registration
* UURIFactory, canonicalize/**, PathologicalPathDecideRule, ExtractorHTML
use recycled matchers via TextUtils
* CrawlURI, KeyedProperties, OverlayContext, SheetOverlaysManager
use ArrayList and indexed access rather than LinkedList and iterator instances
setting "independentExtractors" - when enabled, extractors run regardless of
whether other extractors have run
* profile-crawler-beans.cxml, AbstractFrontier.java, ExtractorParameters.java,
Extractor.java
new setting independentExtractors
* ContentExtractor.java
shouldProcess() - respect independentExtractors