* CheckpointService
Made checkpointIntervalMinutes setter accept a long instead of an int (was missed when the class variable was made a long). While it was non-harmful the way it was, this is more in line with Java conventions.
* AbstractFrontier
remove inbound/outbound queues; simplify managerThread
perform findEligible/schedule/receive/finish immediately
* WorkQueueFrontier
eliminate holdQueues setting
synchronize sendToQueue on target queue
make findEligibleUri reentrant; rely on [Blocking|Stored]Queue thread-safety
make arriving at front of ready trigger for 'active'/budget-session
simplify inactiveQueues management
simplify waking from overflow
* WorkQueue
redefine 'active' as in-budget-session
rename 'held' as 'managed'
synchronize major methods, relying on allQueues cache that operations on same intended queue use identical instance
split over-budget test to isOverSessionBudget and isOverTotalBudget
have toString show classKey
* BdbFrontier
checkpoint fixes: remember nextOrdinal, rotate frontier-recover log, reset queues on recovery
consistencyCheck method for probing state during stress tests
* BdbModule
store lengths in JDB manifest; limit extension of last JDB on recover
* CrawlerJournal
rotate frontier-recover on checkpoints like other logs
* RecordingInputStream.java
check for interrupt on each socket-timeout
* CrawlController.java
on second requestCrawlStop, interrupt threads via ToePool.cleanup
* ToePool.java
adjust for new earlier/repeated cleanup
* ToeThread.java
better warnings/recovery on forced-interrupts
* UriProcessingFormatter.java
allow pre-caching of a specific LogRecord's formatted version, outside the synchronized publish()
* GenerationFileHandler.java
force a format before publish(), to benefit above (a small hit in other cases where it's redundant)
* (many)
distinction between these two constant-collecting classes was fuzzy (to the point that they already referred to each other), and they were both in same subproject/package already as well; so, merged constants into the larger, older class
* WorkQueueFrontier
(findEligibleURI) avoid outbound.capacity-sensitive activation, which under a race created by recent changes led to infinite recursion here
* AbstractFrontier
(next) when nothing is immediately ready, try adding one-at-a-time to outbound, rather than fillOutbound()
* AbstractFrontier
(drainInbound) don't assume current thread will be able to take() full batch; poll() and exit early if other threads have drained queue first
* AbstractFrontier, WorkQueueFrontier
synchronize around all InEvent/findEligibleURI actions, so that drainInbound and fillOutbound may be called outside managerThread
whenever a ToeThread would block on enqueue() or next(), try the appropriate catch-up method before blocking
remove assertions no longer true with activity happening outside managerThread
* UURI
return to custom-serialization rather than externalization
* CachedBdbMap
remove obsolete class
* StripSessionCFIDs, StripSessionIDs
update patterns to restore intended case-insensitivity
* StripSessionCFIDsTest
put expected, actual in expected order
* AutoKryo
extension of Kryo to allow classes to control their own registration, trigger registration of associated classes, and deserialize classes without no-arg constructors
* KryoBinding
binding for use with BDB that uses AutoKryo serialization for a 2X-4X reduction in byte[] size
* UURI
improved serialization via Externalizable and Kryo's CustomSerialization methods
* BdbModule
discard deprecated CachedBDBMap option
(getObjectCache) extend with both declaredClass and valueClass (for when map values are specializations of the declared type, as with frontier.allQueues)
adjust type declarations
* ObjectIdentityBdbCache
use KryoBinding rather than SerialBinding
* CachedBdbMapTest
discarded
* BdbFrontier, BdbServerCache, StatisticsTracker
adjust type declarations, objectCache creation
* BdbMultipleWorkQueues
use KryoBinding rather than (Recycling)SerialBinding
* BdbWorkQueue, CrawlServer, CrawlHost, CrawlURI
add autoregister support
* LinkContext
public for kryo registration
* UURIFactory, canonicalize/**, PathologicalPathDecideRule, ExtractorHTML
use recycled matchers via TextUtils
* CrawlURI, KeyedProperties, OverlayContext, SheetOverlaysManager
use ArrayList and indexed access rather than LinkedList and iterator instances
setting "independentExtractors" - when enabled, extractors run regardless of
whether other extractors have run
* profile-crawler-beans.cxml, AbstractFrontier.java, ExtractorParameters.java,
Extractor.java
new setting independentExtractors
* ContentExtractor.java
shouldProcess() - respect independentExtractors