* KryoBinding.java
wrap ObjectBuffers with WeakReference so that can be garbage collected if necessary (when recreated they'll be back at the default size of 16k)
Reported by Kenji who says, "My guess is that root cause is Recorder.calcRecommendedCharBufferSize():
return Math.min(inStream.getRecordedBufferLength()/2,(int)inStream.getSize());
URL above returns HTML larger than 5GB (infinite smileys!! what the heck), and (int)intStream.getSize() became negative."
* Recorder.java
- return Math.min(inStream.getRecordedBufferLength()/2,(int)inStream.getSize());
+ return (int) Math.min(inStream.getRecordedBufferLength()/2, inStream.getSize());
* WorkQueueFrontier.java
activateInactiveQueue(), deactivateQueue() - synchronize around updates of inactiveQueuesByPrecedence and highestPrecedenceWaiting
* Heritrix.java
instanceMain() - move logging of java vendor/version ahead of version check, since the information could be particularly helpful when version check doesn't pass
* CrawlController.java
requestCrawlResume() - do not create ToePool (creation here appeared to be inherited cruft), paused crawl must already have toe pool
* UriUtils.java
isLikelyFalsePositive() - check for unusual characters, likely mimetypes, and a couple of other common traps
* UriUtilsTest.java
add some tests for this code
Elaboration: CrawlJob needs all beans to receive FINISHED event before teardown. Can't easily guarantee CrawlJob receives the event last. Actually, spring provides a way, org.springframework.core.Ordered, but to make a bean receive events last, *every other bean* must implement Ordered. So this way is simpler.
* CrawlController.java
new member variable isStopComplete and event StopCompleteEvent
* CrawlJob.java
rely on cc.isStopComplete() or StopCompleteEvent to indicate ready for teardown
instantiateContainer() - do not call teardown() on exception from ac.refresh() - can throw IllegalStateException
teardown() - put all uses of variable cc inside not-null check
* PathSharingContext.java
start(), doStart(), stop(), doStop() - remove these overloaded methods, because the spring bug they were working around seems to be fixed, not seeing any problems with cyclical dependencies - might be this one https://jira.springsource.org/browse/SPR-7266
doClose() - remove because superclass version seems to work fine (this method generally wasn't being called anyway, though it would be now with other changes in this checkin)
* CrawlJob.java
refactor teardown to call close() on the ApplicationContext, which calls destroy() on any beans that implement DisposableBean - this is now the way to have beans do stuff at teardown
* CrawlController.java
send FINISHED crawl state event after calling appCtx.stop() so that isFinished() can indicate ready-ness for teardown
* BdbModule.java
move close() to teardown, i.e. implementation of DisposableBean.destroy(); remove shutdown hook and rely on teardown; related tweaks
* WorkQueueFrontier.java, CrawlerLoggerModule.java, BdbUriUniqFilter.java
move close() to teardown
* CrawlMapper.java, AbstractFrontier.java, PreloadedUriPrecedencePolicy.java, FetchWhois.java, FetchHTTP.java, PersistLogProcessor.java, WriterPoolProcessor.java
add comments about cleanup that maybe should wait until teardown
* UriUniqFilter.java
remove incorrect(?) comment
* ToeThread.java
run() - set local variable curi=null when finished with it, because the jvm seems to be holding on to it even after the containing block is finished
* URIAuthorityBasedQueueAssignmentPolicy.java
remove field kp, accessor getKeyedProperties(); instead inherit these from superclass QueueAssignmentPolicy
* BasicProfileTest.java, SelfTestBase.java, CrawlControllerTest.java, PrecedenceLoader.java, MigrateH1to3Tool.java, CrawlerLoggerModule.java, StatisticsTracker.java, CheckpointUtils.java, BdbUriUniqFilter.java, ARCWriterProcessorTest.java, WARCWriterProcessorTest.java, PersistProcessor.java, WriterPoolProcessor.java, PrefixFinderTest.java, StoredQueueTest.java, FileUtilsTest.java, ObjectIdentityBdbManualCacheTest.java, ObjectIdentityBdbCacheTest.java, ObjectPlusFilesOutputStream.java, TestUtils.java, TmpDirTestCase.java, Engine.java
Replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir). In these cases the calls were either already in a spot where the possible IOException would be handled appropriately, or the line was trivially moved into such a block.
* ActionDirectory.java, Engine.java
replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
* BdbModule.java
setup() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException" - conveniently the place where this method is called was already in a try block that catches IOException
* Checkpoint.java
generateFrom() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException"
* CheckpointService.java
move call to Checkpoint.generateFrom() inside existing try block since it now can throw IOException
* Recorder.java
ensure(File) - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
new Recorder(File,String,int,int) - call ensure() on the correct object, the containing directory; and remove redundant call to ensure()
fix mistake in r7240 - 'val.setIdentityCache()' was no longer being called the first time through, when the object comes from the 'supplier' (thanks Gordon)
* WorkQueueFrontier.java
(HER-1925) avoid 2 possible null pointer dereferences
* CrawlController.java
getState() - change declared return type from Object to State
* CrawlerLoggerModule.java
(HER-1932) remove unused, unset field "reports"
* BdbCookieStorage.java
elide pointless extra variable
* FetchHTTP.java, DownloadURLConnection.java, ProcessUtils.java
(HER-1933) use Arrays.toString() for logged arrays
* CrawlServer.java
- remove unused field robotstxtChecksum
- updateRobots() - avoid reinventing existing utility class InstanceofPredicate
* ExternalGeoLookupInterface.java
(HER-1938) extend Serializable, since ExternalGeoLocationDecideRule is declared Serializable and has a ExternalGeoLookupInterface field
* DecideRuleSequence.java
(HER-1937) make field fileLogger transient
* PersistLogProcessor.java
(HER-1924) remove field recoveryCheckpoint shadowing same field in superclass Processor
* ExtractorUniversal.java
(HER-1920) use return value of potentialTLD.toLowerCase() as it appears was intended
* ARCWriterProcessor.java
(HER-1928) avoid possible null pointer dereference
* S3URLConnection.java
(HER-1933) set S3ServiceException as cause of rethrown IOException, and do not put stacktrace array in message
* ObjectIdentityMemCache.java
(HER-1929) avoid possible null pointer dereference
* .classpath
more source jar references, other cleanup
* WorkQueueFrontier.java
findEligibleURI() - propagate queueTotalBudget (also sessionBudget) to ensure the settings are current right before checking if queue is over budget
* DispositionProcessor.java
innerProcess() - call server.updateRobots(curi) after special ignore-robots short-circuiting of retry cycle rather than before, so that updateRobots() deemed-not-found handling applies
* FetchHTTP.java
call curi.setHttpMethod() right after creating the HttpMethod, instead of when processing the response; this is needed so that updateRobots() check of fetchType works properly (before this change, fetchType was "UNKNOWN" on server error)
* PreconditionEnforcer.java, CrawlServer.java
fix typos
* WorkQueueFrontier.java
considerIncluded(CrawlURI) - call FrontierPreparer.prepare(CrawlURI) so that CrawlURI is canonicalized etc before being considered included (similar to what schedule() does)
* CrawlURI.java
getCanonicalString(), getPolitenessDelay() - use logger instead of System.err
testScriptTagWritingScriptType() - As of r7225, ExtractorJS finds a new url - avoid using a strategically placed space. (The fact that UriUtils.isLikelyUri() returns true for the string ");document.write(unescape(" is kind of crazy, but a separate issue)