* FetchHTTP.java
"Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8"
by default - this is what my firefox 4.0 for mac sends
* profile-crawler-beans.cxml
commented-out default value
* ExtractorXML.java
if charset not spec'd in http header look for <?xml encoding=""?>
* ExtractorHTML.java
lookForEncodingInContent() -
1. look for <meta http-equiv="content-type"...>
2. if not found then look for <meta charset="">
3. if not found then <?xml encoding=""...?>
* Recorder.java
setCharacterEncoding() - If new encoding is different from old encoding,
close replayCharSequence and set to null, which will trigger recreation on
next retrieval.
* UriUtils.java
NAIVE_LIKELY_URI_PATTERN - revert change and unpublicize
NAIVE_URI_EXCEPTIONS - add some mimetype strings that came up in testing
isLikelyFalsePositive() - unpublicize
* ExtractorXML.java
XML_URI_EXTRACTOR - do not use NAIVE_LIKELY_URI_PATTERN
shouldExtract() - check for mimetype application/vnd.openxmlformats which is not xml
processXml() - use UriUtils.isLikelyUri()
starts with <?xml..., a@href links are treated as embeds
* ExtractorXML.java
shouldExtract() - return true if content starts with "<?xml" only if it
does not also contain "<!doctype html" or "<html" early in the content
* ExtractorXML.java
instead of considering only strings that start with http(s):, consider all
strings that match UriUtils.NAIVE_LIKELY_URI_PATTERN (and, as before,
constitute the entirety of the xml tag content or attribute value)
* UriUtils.java
- NAIVE_LIKELY_URI_PATTERN - add quotes to the excluded characters so it
does not eat the closing quote when matching xml attribute values, and
make visibility public for use in ExtractorXML
- isLikelyUri() - refactor false positive check into new method
isLikelyFalsePositive() so that it can be used in ExtractorXML avoiding
redundant check against NAIVE_LIKELY_URI_PATTERN
* RobotsDirectives.java
- use plain ConcurrentSkipListSet<String> instead of PrefixSet, to
maintain complete list of prefixes, so we can know the longest match
- allows(String) - return true if longest matching Allow prefix is longer
than or equal to longest matching Disallow prefix
* RobotstxtTest.java
flip expected result of test of generic Allow against specific Disallow
[HER-728] Offer replay stream that has been un-chunked (whether because response was HTTP/1.1 or used chunked in HTTP/1.0 against spec)
[HER-1876] Offer HTTP/1.1 option - for chunked transfer-encoding (but not persistent connections)
* (all)
update to use new (decoded-as-necessary) content streams/CharSequences
[HER-728] Offer replay stream that has been un-chunked (whether because response was HTTP/1.1 or used chunked in HTTP/1.0 against spec)
[HER-1876] Offer HTTP/1.1 option - for chunked transfer-encoding (but not persistent connections)
* FetchHTTP
add 'acceptCompression' and 'useHTTP11' properties, both default false
* Recorder
track whether recorded-input is transfer-encoded (chunked) or content-encoded (gzip etc)
offer alternate replay streams for
(1) raw 'messageBody';
(2) entity (un-chunked if necessary)
(3) content (decompressed if necessary)
always content & GenericReplayCharSequence for CharSequence replays
* GenericReplayCharSequence
always use a stream (rather than random-access buffer)
always decode to prefix buffer first, so short content never touches disk no matter the encoding
* InMemoryReplayCharSequence
deleted; 'Generic' now works similarly for small content and anyway random-access for single-byte-encodings is now rarely possible (given deconding streams)
* RecordingInputStream, RecordingOutputStream
adjust for changed stream names, functionality moved to Recorder
* ReplayCharSequence
use Charset instances rather than names
* ReplayInputStream
add convenience constructor (and tmp-file-destroy) for copying any other inputStream into a seekable ReplayInputStream
Invalid property 'timeoutSeconds' of bean class
[org.archive.modules.fetcher.FetchFTP]: Bean property 'timeoutSeconds' is not
writable or has an invalid setter method. Does the parameter type of the setter
match the return type of the getter?"
* FetchFTP.java
change setTimeoutSeconds(Integer) to setTimeoutSeconds(int)
* FetchFTP.java
setSoTimeoutMs(), getSoTimeoutMs()
(init) - set default value of 20000 ms, same as FetchHTTP
fetch() - set various timeouts on FTPClient - effectively sets connect
timeout and socket timeout on control connection, and socket timeout on
data connection (notably, does not set connect timeout on data connection
w/ current commons-net)
that or preloadSource (a filesystem path), rather than try to handle both
possibilities with one string.
* PersistLoadProcessor.java
* PersistProcessor.java
(regression from h1)
* PersistLoadProcessor.java
change preloadSource from ConfigFile to String, allowing the logic in
PersistProcessor.copyPersistSourceToHistoryMap() to decide how to load it
* URIAuthorityBasedQueueAssignmentPolicy.java
whois urls all go in special "whois" queue
* PreconditionEnforcer.java, CrawlURI.java
move markPrerequisite() method from PreconditionEnforcer to CrawlURI
* profile-crawler-beans.cxml
commented-out whois fetcher clause
* FetchStatusCodes.java
new status codes for whois
* ServerCache.java
getHostFor() - return "whois:" for whois urls, following dns: convention
* UURIFactory.java
ugly hack for whois urls
* SchemeNotInSetDecideRule.java
add whois to list of supported schemes
* WriterPoolProcessor.java
shouldWrite() - we should write successful whois fetches
* WARCWriterProcessor.java
write whois records
* FetchWhois.java
the fetcher
* CookieSpecBase
use Guava InternetDomainName public-suffixes to both (1) prevent storing cookie on public-suffix; (2) prevent looking-up cookies for public-suffixes
* BdbCookieStorage
correct checkpoint/relaunch behavior: don't reuse persisted cookies on normal relaunches; do reuse on resume-from-checkpoint
* WARCWriterProcessor
move checkSize into block protected by catch/finally cleanup of errors
* WARCWriterProcessorTest
test to simulate record-write-failure and verify does not cause pool lockup
* WARCWriterTest, WriterPoolMember
comment and warnings cleanup
* Robotstxt, RobotstxtTest
treat wildcardRules specially, so that it is possible to ask for only explicitly-named matches
* MostFavoredRobotsPolicy
improve comment, change masquerade default to true
* FirstNamedRobotsPolicy, FirstNamedRobotsPolicyTest
policy to consider alternate user-agents in order, using the first that has any explicit (non-wildcard) rules declared (or failing that, the original default user-agent)
* TooManyHopsDecideRule
remove inadvertent reversal of sense of comparison, making default scope reject almost everything (and breaking most self-tests)
* StatisticsTracker
move small stat maps to in-memory ConcurrentMaps
update processedSeedRecords for new ObjectIdentityCache shape
move hostsDistribution/hostsBytes/hostsLastFinished tracking to CrawlHosts/hostsCache
* CrawlSummaryReport
gets hosts count from hostsCache
* SeedRecord
implement IdentityCacheable
* FetchStats
remember last-success times
* IdentityCacheable
new interface required of objects stored in ObjectIdentityCaches
* IdentityCacheableWrapper
wrapper for storing arbitrary objects in ObjectIdentityCaches
* ObjectIdentityCache
keys now always Strings
values now always IdentityCacheables
new dirtyKey() method to ensure a key is persisted
* ObjectIdentityBdbCache, ObjectIdentityMemCache
update for new ObjectIdentityCache shape; still using legacy GC magic
* ObjectIdentityBdbManualCache
alternate implementation relying on dirtying to ensure persistence
use Guava library MapMaker for soft memMap; capped dirtyMap
* BdbModule
use ObjectIdentityBdbManualCache by default
up expectedConcurrency default to 64
* BdbModuleTest, ObjectIdentityBdbCacheTest, ObjectIdentityBdbManualCacheTest
update, add tests
* Frontier
FrontierGroup as IdentityCacheable
* AbstractFrontier, WorkQueueFrontier, BdbFrontier
touchup queue/group accessors
make queue instances dirty whenever mutated
* WorkQueue, CrawlHost, CrawlServer
IdentityCacheable support; appropriate makeDirty()s
* ServerCache, DefaultServerCache
update for new ObjectIdentityCache shape
* CrawlURI
MAX_HOPS_DISPLAYED constant, 50
build new CrawlURI pathFromSeed value via extendHopsPath, which shows only MAX_HOPS_DISPLAYED hops, and precedes string with int value of hops unshown
* CrawlURITest
test for extendHopsPath
* ExtractorJS
apply JS literal string deescaping before interpretation as link
* ExtractorJSTest
3 basic tests, including one for this issue
* Link
provide logging and fallback for case where via may be null; risk identified in creation of test cases could present again if a '.js' file is ever provided as a (via-less) seed
addStats() - use putIfAbsent() to avoid race conditions; move
creation of stats map to constructor to avoid another race condition
* WARCWriter.java
switch back to thread-unsafe data structures for stats since object is used
in one thread at a time according to javadoc
processing is finished.
* Recorder.java
getReplayCharSequence() - remember ReplayCharSequence, obtain new one if
none remembered or remembered has been closed
endReplays() - close replayCharSequence
* ToeThread.java
run() - call Recorder.endReplays() when finished with uri
* ReplayCharSequence.java
isOpen() - return false if close() has been called
* GenericReplayCharSequence.java, InMemoryReplayCharSequence.java
implement isOpen()
* KeyWordProcessor.java, ExtractorHTML.java, HTTPContentDigest.java,
ExtractorCSS.java, ExtractorXML.java
do not close ReplayCharSequence
* Arc2Warc, Warc2Arc
adapt to new constructors, settings-object
* MiserOutputStream
stream to monitor position, optionally suppress flushes
* WriterPoolMember
use MiserOutputStream rather than special file-access to find compressed offsets
use shared settings object rather than copying multiple values
* WriterPoolSettings
add settings for frequentFlushes and writeBufferSize for IO optimization
* ARCReader, ARCWriter, ARCWriterPool
adapt to new constructors, settings-object
* WriterPoolSettingsData
impl of WriterPoolSettings for testing and adhoc use
* WARCWriter
adapt to new constructors, settings-object
remove checkSize rollover to new file (done elsewhere)
remove extraneous id-generator code/static-method
* WARCWriterPool
adapt to new constructors, settings-object
* WARCWriterPoolSettings, WARCWriterPoolSettingsData
add recordIDGenerator setting, impl class for testing/adhoc use
* WARCWriterTest
adapt to new constructors, settings-object
* Generator, GeneratorFactory
removed in favor of RecordIDGenerator interface
* RecordIDGenerator
common interface for classes that can provide WARC record IDs
* UUIDGenerator, UUIDGeneratorTest
adapt to derive from RecordIDGenerator
avoid throwing exceptions
* ARCWriterPoolTest, ARCWriterTest
adapt to new constructors, settings-object
* WARCWriterProcessor
do checkSize file-rollover here, so related records don't span WARCs
calculate processor's totalBytesWritten via position offset changes (as before stats-additions)
delegate record-ID generation
* WriterPoolProcessor
add frequentFlushes and writeBufferSize from new WARCWriterPoolSettings interface
* WriterPoolMember.java
write(*), copyFrom() - return number of (uncompressed) bytes written, for
convenience
* WARCWriter.java
tally by record type: number of records written, content bytes, total
bytes, and (possibly compressed) size on disk
* WARCWriterProcessor.java
keep totals and write report to processors-report.txt
new property textSource, an org.archive.io.ReadSource, allowing
specification inline in cxml as well as in external file
* profile-crawler-beans.cxml
commented out clauses demonstrating use of textSource
* DecideRuleSequence.java
- use Spring Lifecycle start() to initialize logToFile
- inject logger module as SimpleFileLoggerProvider, since we're in
heritrix-modules and don't have access to CrawlerLoggerModule from
heritrix-engine
* CrawlerLoggerModule.java
- implement SimpleFileLoggerProvider
- setupSimpleLog() - use 'T' instead of '+' between date and time in
timestamp
* SimpleFileLoggerProvider.java
new interface with one method, setupSimpleLog()
* Scoper.java
do not call scope.start()/scope.stop(), these are handled using Spring
Lifecycle now
* DecideRule.java
remove unused, unneeded start()/stop()