* URIAuthorityBasedQueueAssignmentPolicy.java
remove field kp, accessor getKeyedProperties(); instead inherit these from superclass QueueAssignmentPolicy
* BasicProfileTest.java, SelfTestBase.java, CrawlControllerTest.java, PrecedenceLoader.java, MigrateH1to3Tool.java, CrawlerLoggerModule.java, StatisticsTracker.java, CheckpointUtils.java, BdbUriUniqFilter.java, ARCWriterProcessorTest.java, WARCWriterProcessorTest.java, PersistProcessor.java, WriterPoolProcessor.java, PrefixFinderTest.java, StoredQueueTest.java, FileUtilsTest.java, ObjectIdentityBdbManualCacheTest.java, ObjectIdentityBdbCacheTest.java, ObjectPlusFilesOutputStream.java, TestUtils.java, TmpDirTestCase.java, Engine.java
Replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir). In these cases the calls were either already in a spot where the possible IOException would be handled appropriately, or the line was trivially moved into such a block.
* ActionDirectory.java, Engine.java
replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
* BdbModule.java
setup() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException" - conveniently the place where this method is called was already in a try block that catches IOException
* Checkpoint.java
generateFrom() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException"
* CheckpointService.java
move call to Checkpoint.generateFrom() inside existing try block since it now can throw IOException
* Recorder.java
ensure(File) - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
new Recorder(File,String,int,int) - call ensure() on the correct object, the containing directory; and remove redundant call to ensure()
* WorkQueueFrontier.java
(HER-1925) avoid 2 possible null pointer dereferences
* CrawlController.java
getState() - change declared return type from Object to State
* CrawlerLoggerModule.java
(HER-1932) remove unused, unset field "reports"
* BdbCookieStorage.java
elide pointless extra variable
* FetchHTTP.java, DownloadURLConnection.java, ProcessUtils.java
(HER-1933) use Arrays.toString() for logged arrays
* CrawlServer.java
- remove unused field robotstxtChecksum
- updateRobots() - avoid reinventing existing utility class InstanceofPredicate
* ExternalGeoLookupInterface.java
(HER-1938) extend Serializable, since ExternalGeoLocationDecideRule is declared Serializable and has a ExternalGeoLookupInterface field
* DecideRuleSequence.java
(HER-1937) make field fileLogger transient
* PersistLogProcessor.java
(HER-1924) remove field recoveryCheckpoint shadowing same field in superclass Processor
* ExtractorUniversal.java
(HER-1920) use return value of potentialTLD.toLowerCase() as it appears was intended
* ARCWriterProcessor.java
(HER-1928) avoid possible null pointer dereference
* S3URLConnection.java
(HER-1933) set S3ServiceException as cause of rethrown IOException, and do not put stacktrace array in message
* ObjectIdentityMemCache.java
(HER-1929) avoid possible null pointer dereference
* .classpath
more source jar references, other cleanup
* WorkQueueFrontier.java
findEligibleURI() - propagate queueTotalBudget (also sessionBudget) to ensure the settings are current right before checking if queue is over budget
* DispositionProcessor.java
innerProcess() - call server.updateRobots(curi) after special ignore-robots short-circuiting of retry cycle rather than before, so that updateRobots() deemed-not-found handling applies
* FetchHTTP.java
call curi.setHttpMethod() right after creating the HttpMethod, instead of when processing the response; this is needed so that updateRobots() check of fetchType works properly (before this change, fetchType was "UNKNOWN" on server error)
* PreconditionEnforcer.java, CrawlServer.java
fix typos
* WorkQueueFrontier.java
considerIncluded(CrawlURI) - call FrontierPreparer.prepare(CrawlURI) so that CrawlURI is canonicalized etc before being considered included (similar to what schedule() does)
* CrawlURI.java
getCanonicalString(), getPolitenessDelay() - use logger instead of System.err
add 'trackSources' to make hosts-by-source-tag report optional in large crawls, where it is very expensive (and an OOME risk until other changes are also made)
* Report.java
make StatisticsTracker stats an argument to write() instead of a field; new properties shouldReportAtEndOfCrawl and shouldReportDuringCrawl, both default true, but configurable if desired, replacing the LIVE_REPORTS/END_REPORTS stuff in StatisticsTracker
* StatisticsTracker.java
new property "reports", a List<Report>, which defaults (with lazy initialization) to the same reports we've always had, replacing hard-coded lists of reports classes and related code
* profile-crawler-beans.cxml
default reports list, commented out
* SeedsReport.java FrontierSummaryReport.java MimetypesReport.java SourceTagsReport.java FrontierNonemptyReport.java CrawlSummaryReport.java ResponseCodeReport.java HostsReport.java ProcessorsReport.java ToeThreadsReport.java
update write(PrintWriter, StatisticsTracker) arguments
* JobResource.java
replace use of StatisticsTracker.LIVE_REPORTS
contributed by Adam Wilmer
* FetchHTTP.java
properties, setup for http proxy user/password
* profile-crawler-beans.cxml
commented-out example settings for proxy auth properties
* ActionDirectory.java, CrawlerLoggerModule.java, StatisticsTracker.java, SurtPrefixedDecideRule.java, WriterPoolProcessor.java, profile-crawler-beans.cxml
use camelcase ${launchId}
* CheckpointService.java, JobResource.java
rename getAvailableCheckpointDirectories() to findAvailableCheckpointDirectories() so that it's not treated as a bean property
* WriterPoolProcessor.java, WriterPoolSettingsData.java, WriterPoolMember.java
rename getOutputDirs() to calcOutputDirs() so that it's not treated as a bean property
* ConfigPath.java
let ConfigPathConfigurer to do interpolation of ${launchId}
* ConfigFile.java
let ConfigPathConfigurer do the snapshotting of config file
* CrawlJob.java, PathSharingContext.java
remove launchDir initialization out of CrawlJob into PathSharingContext
* ConfigPathConfigurer.java
- remove special handling of WriterPoolProcessor store paths; instead, look in beans for ConfigPaths within Iterables
- have each ConfigPath hold reference to this ConfigPathConfigurer to use for interpolating ${launchId} and snapshotting config files
* CrawlJob.java
decided to go with plain timestamp17 as launch id, no "launch-" prefix
* ConfigPathConfigurer.java
refactored special remembering of WriterPoolProcessor storePaths out of fixupPaths()
* UriProcessingFormatter, Preformatter, GenerationFileHandler
improve preformat-outside-synchronized optimization so that the LogRecord/CrawlURI doesn't linger until next displaces it
* CrawlURI
(processingCleanup) null more of last-processing-run collected values
* HardLinker.java
renamed FilesystemLinkMaker.java
* FilesystemLinkMaker.java
add support for symbolic links
* CLibrary.java
new method symlink()
* BdbModule.java
use new class name FilesystemLinkMaker
* CrawlJob.java
at crawl launch, create launch directory launch-{timestamp17}, copy cxml there, symlink "current" to launch dir, inform ConfigPaths
* ConfigPath.java
interpolate ${launch-id} in configured paths
* ConfigFile.java
obtainReader() - snapshot config files to launch dir when they are read
* ActionDirectory.java
default doneDir now ${launch-id}/actions-done
actOn() - symlink from old style done dir action/done to done files
* SurtPrefixedDecideRule.java
default surtsDumpFile now ${launch-id}/surts.dump
pathsFixedUp() - this gets called at build time, but we don't want anything written to disk until launch time, so remove call to dumpSurtPrefixSet() here
* CrawlerLoggerModule.java
default logs dir now ${launch-id}/logs
* StatisticsTracker.java
default reports dir now ${launch-id}/reports
* WriterPoolProcessor.java
default writer base path now ${launch-id}
* profile-crawler-beans.cxml
update to reflect new default paths under launch dirs
* PropertyUtils.java
fix javadoc typo
* pom.xml, .classpath
update references to necessary spring-3.0.5 packages
* SeedModule
merge rather than replace event listeners, so that (now later) autowiring doesn't clobber the self-insertion done by anonymous (non-top-level) SurtPrefixedDecideRule bean
* HeritrixLifecycleProcessor, PathScharingContext
use this new non-default LifecycleProcessor to avoid new Spring3 behavior of auto-start()ing context on refresh(build)
* CrawlController
example of using @Value annotation to set default value: will offer benefits for auto-discovery of defaults for configuration interface, or enforcing maximally-explicit configurations (see [HER-1897])
* profile-crawler-beans.cxml, selftest-crawler-beans.cxml
update templates with spring3 preamble/boilerplate
[HER-1910] H3: job-related pages very slow to render (pause mid-way through) in large crawls
* BdbModule
option to reuse bdb data on StoredQueue-creation
* Checkpoint
new saveWriter, loadReader utilities for creating extra checkpoint files (for the long list of ready queues)
* BdbFrontier
on checkpoint, remember ordered list of inProcess/ready/snoozed queues (so same queues are active upon resume)
on resume, reuse old retired/inactive StoredQueue data, and reload active queues from newly-saved list
* StoredQueue
restore tailIndex properly when resuing old data
more-efficient size() calculation (for HER-1910)
* JobRelatedResource
catch InvalidPropertyException, render as red error message rather than fouling entire trace of beans to web UI
* StatisticsTracker
rename getSnapshots to listSnapshots to prevent interpretation as (invalid) bean-property
* TextSeedModule
refactor announceSeeds to occur in background thread if non-default blockAwaitingSeedLines value is set
signal CountDownLatch on each line, allowing calling thread to proceed at right count
* profile-crawler-beans.cxml
commented-out blockAwaitingSeedLines default settings (-1, meaning wait for all seed lines)
* CrawlController
change pauseAtFinish to runWhileEmpty
add EMPTY state
replace isStateRunning() with isActive() (RUNNING or EMPTY)
trigger finish-test off frontier EMPTY rather than PAUSE
in logging say 'running' rather than 'resumed'
* Frontier
add EMPTY state
* AbstractFrontier
handle EMPTY like RUNNING, transition EMPTY<->RUNNING when appropriate
* StatisticsTracker
handle EMPTY analogous to RUNNING
in logging say 'running' rather than 'resumed'
* CrawlJob, DiskSpaceMonitor
replace isStateRunning() with isActive()
redirect url sometimes '0 NOTCRAWLED' in seeds-report.txt even when crawled"
* CandidatesProcessor.java
innerProcess() - set force-fetch on outlinks promoted to seeds
* CrawlURI.java
new field JSONObject extraInfo, methods getExtraInfo() and addExtraInfo()
* WARCWriterProcessor.java
write() - on success, add warcFilename to CrawlURI extraInfo
* CrawlerLoggerModule.java
new config option logExtraInfo, default false
* UriProcessingFormatter.java
new field logExtraInfo
format() - include CrawlURI extraInfo if logExtraInfo is enabled
also include "-" if CrawlURI has no annotations, since this is no longer the last field on the line
* profile-crawler-beans.cxml
<!-- <property name="logExtraInfo" value="false" /> -->
* NonFatalErrorFormatter.java, RuntimeErrorFormatter.java
new constructor to handle logExtraInfo since these subclass UriProcessingFormatter
even when crawled, when original seed also has a regular link to the redirect
url
* CandidatesProcessor.java
innerProcess() - present seed outlinks to the frontier ahead of non-seed outlinks, so that seed version of any duplicated outlink is always the one that's crawled