* CrawlController.java
requestCrawlResume() - do not create ToePool (creation here appeared to be inherited cruft), paused crawl must already have toe pool
* UriUtils.java
isLikelyFalsePositive() - check for unusual characters, likely mimetypes, and a couple of other common traps
* UriUtilsTest.java
add some tests for this code
Elaboration: CrawlJob needs all beans to receive FINISHED event before teardown. Can't easily guarantee CrawlJob receives the event last. Actually, spring provides a way, org.springframework.core.Ordered, but to make a bean receive events last, *every other bean* must implement Ordered. So this way is simpler.
* CrawlController.java
new member variable isStopComplete and event StopCompleteEvent
* CrawlJob.java
rely on cc.isStopComplete() or StopCompleteEvent to indicate ready for teardown
instantiateContainer() - do not call teardown() on exception from ac.refresh() - can throw IllegalStateException
teardown() - put all uses of variable cc inside not-null check
* PathSharingContext.java
start(), doStart(), stop(), doStop() - remove these overloaded methods, because the spring bug they were working around seems to be fixed, not seeing any problems with cyclical dependencies - might be this one https://jira.springsource.org/browse/SPR-7266
doClose() - remove because superclass version seems to work fine (this method generally wasn't being called anyway, though it would be now with other changes in this checkin)
* CrawlJob.java
refactor teardown to call close() on the ApplicationContext, which calls destroy() on any beans that implement DisposableBean - this is now the way to have beans do stuff at teardown
* CrawlController.java
send FINISHED crawl state event after calling appCtx.stop() so that isFinished() can indicate ready-ness for teardown
* BdbModule.java
move close() to teardown, i.e. implementation of DisposableBean.destroy(); remove shutdown hook and rely on teardown; related tweaks
* WorkQueueFrontier.java, CrawlerLoggerModule.java, BdbUriUniqFilter.java
move close() to teardown
* CrawlMapper.java, AbstractFrontier.java, PreloadedUriPrecedencePolicy.java, FetchWhois.java, FetchHTTP.java, PersistLogProcessor.java, WriterPoolProcessor.java
add comments about cleanup that maybe should wait until teardown
* UriUniqFilter.java
remove incorrect(?) comment
* ToeThread.java
run() - set local variable curi=null when finished with it, because the jvm seems to be holding on to it even after the containing block is finished
* URIAuthorityBasedQueueAssignmentPolicy.java
remove field kp, accessor getKeyedProperties(); instead inherit these from superclass QueueAssignmentPolicy
* BasicProfileTest.java, SelfTestBase.java, CrawlControllerTest.java, PrecedenceLoader.java, MigrateH1to3Tool.java, CrawlerLoggerModule.java, StatisticsTracker.java, CheckpointUtils.java, BdbUriUniqFilter.java, ARCWriterProcessorTest.java, WARCWriterProcessorTest.java, PersistProcessor.java, WriterPoolProcessor.java, PrefixFinderTest.java, StoredQueueTest.java, FileUtilsTest.java, ObjectIdentityBdbManualCacheTest.java, ObjectIdentityBdbCacheTest.java, ObjectPlusFilesOutputStream.java, TestUtils.java, TmpDirTestCase.java, Engine.java
Replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir). In these cases the calls were either already in a spot where the possible IOException would be handled appropriately, or the line was trivially moved into such a block.
* ActionDirectory.java, Engine.java
replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
* BdbModule.java
setup() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException" - conveniently the place where this method is called was already in a try block that catches IOException
* Checkpoint.java
generateFrom() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException"
* CheckpointService.java
move call to Checkpoint.generateFrom() inside existing try block since it now can throw IOException
* Recorder.java
ensure(File) - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
new Recorder(File,String,int,int) - call ensure() on the correct object, the containing directory; and remove redundant call to ensure()
fix mistake in r7240 - 'val.setIdentityCache()' was no longer being called the first time through, when the object comes from the 'supplier' (thanks Gordon)
* WorkQueueFrontier.java
(HER-1925) avoid 2 possible null pointer dereferences
* CrawlController.java
getState() - change declared return type from Object to State
* CrawlerLoggerModule.java
(HER-1932) remove unused, unset field "reports"
* BdbCookieStorage.java
elide pointless extra variable
* FetchHTTP.java, DownloadURLConnection.java, ProcessUtils.java
(HER-1933) use Arrays.toString() for logged arrays
* CrawlServer.java
- remove unused field robotstxtChecksum
- updateRobots() - avoid reinventing existing utility class InstanceofPredicate
* ExternalGeoLookupInterface.java
(HER-1938) extend Serializable, since ExternalGeoLocationDecideRule is declared Serializable and has a ExternalGeoLookupInterface field
* DecideRuleSequence.java
(HER-1937) make field fileLogger transient
* PersistLogProcessor.java
(HER-1924) remove field recoveryCheckpoint shadowing same field in superclass Processor
* ExtractorUniversal.java
(HER-1920) use return value of potentialTLD.toLowerCase() as it appears was intended
* ARCWriterProcessor.java
(HER-1928) avoid possible null pointer dereference
* S3URLConnection.java
(HER-1933) set S3ServiceException as cause of rethrown IOException, and do not put stacktrace array in message
* ObjectIdentityMemCache.java
(HER-1929) avoid possible null pointer dereference
* .classpath
more source jar references, other cleanup
* WorkQueueFrontier.java
findEligibleURI() - propagate queueTotalBudget (also sessionBudget) to ensure the settings are current right before checking if queue is over budget
* DispositionProcessor.java
innerProcess() - call server.updateRobots(curi) after special ignore-robots short-circuiting of retry cycle rather than before, so that updateRobots() deemed-not-found handling applies
* FetchHTTP.java
call curi.setHttpMethod() right after creating the HttpMethod, instead of when processing the response; this is needed so that updateRobots() check of fetchType works properly (before this change, fetchType was "UNKNOWN" on server error)
* PreconditionEnforcer.java, CrawlServer.java
fix typos
* WorkQueueFrontier.java
considerIncluded(CrawlURI) - call FrontierPreparer.prepare(CrawlURI) so that CrawlURI is canonicalized etc before being considered included (similar to what schedule() does)
* CrawlURI.java
getCanonicalString(), getPolitenessDelay() - use logger instead of System.err
testScriptTagWritingScriptType() - As of r7225, ExtractorJS finds a new url - avoid using a strategically placed space. (The fact that UriUtils.isLikelyUri() returns true for the string ");document.write(unescape(" is kind of crazy, but a separate issue)
* ExtractorJS.java
considerStrings() - backtrack to just ahead of previous closing quote to reconsider it as a possible opening quote
* ExtractorJSTest.java
minimal test case
* AbstractCookieStorage.java
make cookiesLoadFile a ConfigFile and use obtainReader() so that the file is snapshotted into launch dir
saveCookies() - include expiration date, so file format match the format it claims to use, and what loadCookies() expects
loadCookies() - handle expiration date field, and actually put each cookie in the cookie map
refactor loadCookies() for use with Reader from ConfigFile.obtainReader()
fix up javadocs some
* CookieSpecBase.java
match(String, int, String, boolean, SortedMap) - use InternetDomainName.name() instead of .toString(), since the latter doesn't return the plain old domain name
* Cookie.java
javadoc typo