* SheetOverlaysManager.java
getOverlayMap(String) - return null if sheet missing instead of triggering npe
* KeyedProperties.java
get(String) - check for null return value from getOverlayMap() and log warning
* KryoBinding.java
wrap ObjectBuffers with WeakReference so that can be garbage collected if necessary (when recreated they'll be back at the default size of 16k)
Reported by Kenji who says, "My guess is that root cause is Recorder.calcRecommendedCharBufferSize():
return Math.min(inStream.getRecordedBufferLength()/2,(int)inStream.getSize());
URL above returns HTML larger than 5GB (infinite smileys!! what the heck), and (int)intStream.getSize() became negative."
* Recorder.java
- return Math.min(inStream.getRecordedBufferLength()/2,(int)inStream.getSize());
+ return (int) Math.min(inStream.getRecordedBufferLength()/2, inStream.getSize());
* UriUtils.java
isLikelyFalsePositive() - check for unusual characters, likely mimetypes, and a couple of other common traps
* UriUtilsTest.java
add some tests for this code
* PathSharingContext.java
start(), doStart(), stop(), doStop() - remove these overloaded methods, because the spring bug they were working around seems to be fixed, not seeing any problems with cyclical dependencies - might be this one https://jira.springsource.org/browse/SPR-7266
doClose() - remove because superclass version seems to work fine (this method generally wasn't being called anyway, though it would be now with other changes in this checkin)
* CrawlJob.java
refactor teardown to call close() on the ApplicationContext, which calls destroy() on any beans that implement DisposableBean - this is now the way to have beans do stuff at teardown
* CrawlController.java
send FINISHED crawl state event after calling appCtx.stop() so that isFinished() can indicate ready-ness for teardown
* BdbModule.java
move close() to teardown, i.e. implementation of DisposableBean.destroy(); remove shutdown hook and rely on teardown; related tweaks
* WorkQueueFrontier.java, CrawlerLoggerModule.java, BdbUriUniqFilter.java
move close() to teardown
* CrawlMapper.java, AbstractFrontier.java, PreloadedUriPrecedencePolicy.java, FetchWhois.java, FetchHTTP.java, PersistLogProcessor.java, WriterPoolProcessor.java
add comments about cleanup that maybe should wait until teardown
* UriUniqFilter.java
remove incorrect(?) comment
* BasicProfileTest.java, SelfTestBase.java, CrawlControllerTest.java, PrecedenceLoader.java, MigrateH1to3Tool.java, CrawlerLoggerModule.java, StatisticsTracker.java, CheckpointUtils.java, BdbUriUniqFilter.java, ARCWriterProcessorTest.java, WARCWriterProcessorTest.java, PersistProcessor.java, WriterPoolProcessor.java, PrefixFinderTest.java, StoredQueueTest.java, FileUtilsTest.java, ObjectIdentityBdbManualCacheTest.java, ObjectIdentityBdbCacheTest.java, ObjectPlusFilesOutputStream.java, TestUtils.java, TmpDirTestCase.java, Engine.java
Replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir). In these cases the calls were either already in a spot where the possible IOException would be handled appropriately, or the line was trivially moved into such a block.
* ActionDirectory.java, Engine.java
replace calls to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
* BdbModule.java
setup() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException" - conveniently the place where this method is called was already in a try block that catches IOException
* Checkpoint.java
generateFrom() - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir) and add "throws IOException"
* CheckpointService.java
move call to Checkpoint.generateFrom() inside existing try block since it now can throw IOException
* Recorder.java
ensure(File) - replace call to File.mkdirs() with FileUtils.ensureWriteableDirectory(dir), and throw IllegalStateException on failure
new Recorder(File,String,int,int) - call ensure() on the correct object, the containing directory; and remove redundant call to ensure()
fix mistake in r7240 - 'val.setIdentityCache()' was no longer being called the first time through, when the object comes from the 'supplier' (thanks Gordon)
* WorkQueueFrontier.java
(HER-1925) avoid 2 possible null pointer dereferences
* CrawlController.java
getState() - change declared return type from Object to State
* CrawlerLoggerModule.java
(HER-1932) remove unused, unset field "reports"
* BdbCookieStorage.java
elide pointless extra variable
* FetchHTTP.java, DownloadURLConnection.java, ProcessUtils.java
(HER-1933) use Arrays.toString() for logged arrays
* CrawlServer.java
- remove unused field robotstxtChecksum
- updateRobots() - avoid reinventing existing utility class InstanceofPredicate
* ExternalGeoLookupInterface.java
(HER-1938) extend Serializable, since ExternalGeoLocationDecideRule is declared Serializable and has a ExternalGeoLookupInterface field
* DecideRuleSequence.java
(HER-1937) make field fileLogger transient
* PersistLogProcessor.java
(HER-1924) remove field recoveryCheckpoint shadowing same field in superclass Processor
* ExtractorUniversal.java
(HER-1920) use return value of potentialTLD.toLowerCase() as it appears was intended
* ARCWriterProcessor.java
(HER-1928) avoid possible null pointer dereference
* S3URLConnection.java
(HER-1933) set S3ServiceException as cause of rethrown IOException, and do not put stacktrace array in message
* ObjectIdentityMemCache.java
(HER-1929) avoid possible null pointer dereference
* .classpath
more source jar references, other cleanup
* CookieSpecBase.java
match(String, int, String, boolean, SortedMap) - use InternetDomainName.name() instead of .toString(), since the latter doesn't return the plain old domain name
* Cookie.java
javadoc typo
* ActionDirectory.java, CrawlerLoggerModule.java, StatisticsTracker.java, SurtPrefixedDecideRule.java, WriterPoolProcessor.java, profile-crawler-beans.cxml
use camelcase ${launchId}
* CheckpointService.java, JobResource.java
rename getAvailableCheckpointDirectories() to findAvailableCheckpointDirectories() so that it's not treated as a bean property
* WriterPoolProcessor.java, WriterPoolSettingsData.java, WriterPoolMember.java
rename getOutputDirs() to calcOutputDirs() so that it's not treated as a bean property
* ConfigPath.java
let ConfigPathConfigurer to do interpolation of ${launchId}
* ConfigFile.java
let ConfigPathConfigurer do the snapshotting of config file
* CrawlJob.java, PathSharingContext.java
remove launchDir initialization out of CrawlJob into PathSharingContext
* ConfigPathConfigurer.java
- remove special handling of WriterPoolProcessor store paths; instead, look in beans for ConfigPaths within Iterables
- have each ConfigPath hold reference to this ConfigPathConfigurer to use for interpolating ${launchId} and snapshotting config files
* Recorder
object to unsupported content-encodings when first set
* FetchHTTP
note as annotation unsupported content-encodings
* Link
not as annotation when base URI is used for absent via
* ExtractorHTML
downgrade char-sequence reading problem (often a chunking problem) to WARNING from SEVERE
* CrawlJob.java
decided to go with plain timestamp17 as launch id, no "launch-" prefix
* ConfigPathConfigurer.java
refactored special remembering of WriterPoolProcessor storePaths out of fixupPaths()
* UriProcessingFormatter, Preformatter, GenerationFileHandler
improve preformat-outside-synchronized optimization so that the LogRecord/CrawlURI doesn't linger until next displaces it
* CrawlURI
(processingCleanup) null more of last-processing-run collected values
* WriterPoolProcessor.java, ARCWriterProcessor.java, WARCWriterProcessor.java
change type of storePaths to List<ConfigPath> and handle appropriately
* ConfigPathConfigurer.java
fixupPaths() - old code did not touch WriterPoolProcessor storePaths, since they're deeply nested inside the bean, but they need to be remembered for later interpolation of ${launch-id}, so add special handling
* HardLinker.java
renamed FilesystemLinkMaker.java
* FilesystemLinkMaker.java
add support for symbolic links
* CLibrary.java
new method symlink()
* BdbModule.java
use new class name FilesystemLinkMaker
* CrawlJob.java
at crawl launch, create launch directory launch-{timestamp17}, copy cxml there, symlink "current" to launch dir, inform ConfigPaths
* ConfigPath.java
interpolate ${launch-id} in configured paths
* ConfigFile.java
obtainReader() - snapshot config files to launch dir when they are read
* ActionDirectory.java
default doneDir now ${launch-id}/actions-done
actOn() - symlink from old style done dir action/done to done files
* SurtPrefixedDecideRule.java
default surtsDumpFile now ${launch-id}/surts.dump
pathsFixedUp() - this gets called at build time, but we don't want anything written to disk until launch time, so remove call to dumpSurtPrefixSet() here
* CrawlerLoggerModule.java
default logs dir now ${launch-id}/logs
* StatisticsTracker.java
default reports dir now ${launch-id}/reports
* WriterPoolProcessor.java
default writer base path now ${launch-id}
* profile-crawler-beans.cxml
update to reflect new default paths under launch dirs
* PropertyUtils.java
fix javadoc typo
* WriterPool
make new LARGEST_MAX_ACTIVE (255) the capacity of the reuse-pool, so that the maxActive property may vary up to that and still work; also handle more gracefully the failure of trying to use a larger number than the capacity (writer gets closed early rather than left hanging open)
* pom.xml, .classpath
update references to necessary spring-3.0.5 packages
* SeedModule
merge rather than replace event listeners, so that (now later) autowiring doesn't clobber the self-insertion done by anonymous (non-top-level) SurtPrefixedDecideRule bean
* HeritrixLifecycleProcessor, PathScharingContext
use this new non-default LifecycleProcessor to avoid new Spring3 behavior of auto-start()ing context on refresh(build)
* CrawlController
example of using @Value annotation to set default value: will offer benefits for auto-discovery of defaults for configuration interface, or enforcing maximally-explicit configurations (see [HER-1897])
* profile-crawler-beans.cxml, selftest-crawler-beans.cxml
update templates with spring3 preamble/boilerplate
[HER-1910] H3: job-related pages very slow to render (pause mid-way through) in large crawls
* BdbModule
option to reuse bdb data on StoredQueue-creation
* Checkpoint
new saveWriter, loadReader utilities for creating extra checkpoint files (for the long list of ready queues)
* BdbFrontier
on checkpoint, remember ordered list of inProcess/ready/snoozed queues (so same queues are active upon resume)
on resume, reuse old retired/inactive StoredQueue data, and reload active queues from newly-saved list
* StoredQueue
restore tailIndex properly when resuing old data
more-efficient size() calculation (for HER-1910)
* HardLinker.java
makeHardLink() - wrap unix link() and windows CreateHardLink() and choose depending on platform
* BdbModule.java
doCheckpoint() - use HardLinker.makeHardLink()
* GenericReplayCharSequence.java
decode() - read another character to check for more content, rather than consult BufferedReader.ready(), since the latter sometimes returns false even when there is more to read
* ReplayCharSequence, GenericReplayCharSequence
use Charset instance rather than String-name
* Recorder
use Charset instance ratehr than String-name
(getContentReplayCharSequence) avoid reusing cached ReplayCharSequence when encoding-in-use has changed since it was created
(getContentReplayPrefixString) allow requested specific charset-interpretation
* ExtractorHTML, ExtractorXML
work with Charset instances
double-check that in-content-declarations are self-consistent before using
recycle matcher instances
change minor charset problems/decisions to a crawl.log annotation
* ExtractorXML.java
if charset not spec'd in http header look for <?xml encoding=""?>
* ExtractorHTML.java
lookForEncodingInContent() -
1. look for <meta http-equiv="content-type"...>
2. if not found then look for <meta charset="">
3. if not found then <?xml encoding=""...?>
* Recorder.java
setCharacterEncoding() - If new encoding is different from old encoding,
close replayCharSequence and set to null, which will trigger recreation on
next retrieval.