Commit Graph
230 Commits
Author SHA1 Message Date
gojomo 39091dc82c Fix tests
* TextSeedModule
    remove unintended @Required annotation
2011-06-22 23:09:43 +00:00
gojomo 842f495ee9 [HER-1905] H3: allow crawling to begin while large seed list still loading
* TextSeedModule
    refactor announceSeeds to occur in background thread if non-default blockAwaitingSeedLines value is set
    signal CountDownLatch on each line, allowing calling thread to proceed at right count
* profile-crawler-beans.cxml
    commented-out blockAwaitingSeedLines default settings (-1, meaning wait for all seed lines)
2011-06-22 21:21:26 +00:00
gojomo c4636f38ec [HER-741] Make extractors interrogate for charset
* FetchHTTP
    use Charset-instances rather than String names
2011-06-04 00:25:23 +00:00
gojomo 2043d1801a [HER-741] Make extractors interrogate for charset
* ReplayCharSequence, GenericReplayCharSequence
    use Charset instance rather than String-name
* Recorder
    use Charset instance ratehr than String-name
    (getContentReplayCharSequence) avoid reusing cached ReplayCharSequence when encoding-in-use has changed since it was created
    (getContentReplayPrefixString) allow requested specific charset-interpretation
* ExtractorHTML, ExtractorXML
    work with Charset instances
    double-check that in-content-declarations are self-consistent before using
    recycle matcher instances
    change minor charset problems/decisions to a crawl.log annotation
2011-06-03 21:28:03 +00:00
gojomo 572aa83d39 * CrawlURI
make extraInfo declarations adjacent
    utility test for charset-declared-in-contentType
2011-06-03 21:17:32 +00:00
gojomo 292b30af1b remove unnecessary Serializable interface 2011-05-28 00:02:38 +00:00
gojomo 083e75e30a * CrawlServer.java
autoregister ConcurrentSkipListSet
2011-05-28 00:02:10 +00:00
nlevitt 728951d1ff Fix for HER-1791 - heritrix writes revisit record whether or not previous fetch
was archived
* CoreAttributeConstants.java
    new flag A_HISTORY_GOOD_TO_STORE w/ javadoc
* WriterPoolProcessor.java
    shouldWrite() - set A_HISTORY_GOOD_TO_STORE on CrawlURI if we decide to skip writing because of previous fetch with identical digest
* WARCWriterProcessor.java
    write() - set A_HISTORY_GOOD_TO_STORE on CrawlURI on successful write to warc
* PersistProcessor.java
    shouldStore() - change to return true if A_HISTORY_GOOD_TO_STORE is set on the CrawlURI
2011-05-26 16:44:53 +00:00
nlevitt 2d267cf173 HER-843 - crawl.log record of which ARC a capture landed in
* WARCWriterProcessor.java
    write() - on success, add arcFilename to CrawlURI extraInfo
2011-05-25 22:23:53 +00:00
nlevitt 5e1571f8dd HER-1790 - record warc where url was saved in crawl.log
* CrawlURI.java
    new field JSONObject extraInfo, methods getExtraInfo() and addExtraInfo()
* WARCWriterProcessor.java
    write() - on success, add warcFilename to CrawlURI extraInfo
* CrawlerLoggerModule.java
    new config option logExtraInfo, default false 
* UriProcessingFormatter.java
    new field logExtraInfo
    format() - include CrawlURI extraInfo if logExtraInfo is enabled
               also include "-" if CrawlURI has no annotations, since this is no longer the last field on the line
* profile-crawler-beans.cxml
    <!-- <property name="logExtraInfo" value="false" /> -->
* NonFatalErrorFormatter.java, RuntimeErrorFormatter.java
    new constructor to handle logExtraInfo since these subclass UriProcessingFormatter
2011-05-25 22:11:18 +00:00
nlevitt e70f0f8a50 Fix HER-1889 should recover from problems unescaping javascript in ExtractorJS
* ExtractorJS.java
    considerStrings() - catch exceptions from
    StringEscapeUtils.unescapeJavaScript(), log warning, and proceed
2011-05-20 00:05:09 +00:00
nlevitt d26bee27e8 * Link.java
addRelativeToVia() - really use base instead of via when via is missing, as
    the warning warns will happen
2011-05-19 23:43:16 +00:00
nlevitt 26f34c1b8c HER-1547 some sites may require an 'Accept' header
* FetchHTTP.java
    "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8"
    by default - this is what my firefox 4.0 for mac sends
* profile-crawler-beans.cxml
    commented-out default value
2011-04-25 20:43:17 +00:00
nlevitt d159b4e106 HER-741 Make extractors interrogate for charset
* ExtractorXML.java
    if charset not spec'd in http header look for <?xml encoding=""?>
* ExtractorHTML.java 
    lookForEncodingInContent() - 
    1. look for <meta http-equiv="content-type"...>
    2. if not found then look for <meta charset="">
    3. if not found then <?xml encoding=""...?>
* Recorder.java
    setCharacterEncoding() - If new encoding is different from old encoding,
    close replayCharSequence and set to null, which will trigger recreation on
    next retrieval.
2011-04-22 18:43:18 +00:00
nlevitt c9c74b5c28 * ExtractorXML.java
shouldExtract() - use getContentReplayPrefixString() for <?xml check
2011-04-22 00:42:09 +00:00
nlevitt a2eaa9d6ac HER-1820 followup
* UriUtils.java
    NAIVE_LIKELY_URI_PATTERN - revert change and unpublicize
    NAIVE_URI_EXCEPTIONS - add some mimetype strings that came up in testing
    isLikelyFalsePositive() - unpublicize
* ExtractorXML.java
    XML_URI_EXTRACTOR - do not use NAIVE_LIKELY_URI_PATTERN
    shouldExtract() - check for mimetype application/vnd.openxmlformats which is not xml
    processXml() - use UriUtils.isLikelyUri()
2011-04-21 17:18:01 +00:00
nlevitt 43dae4d84d HER-1884 when html is processed by ExtractorXML, which can happen when it
starts with <?xml..., a@href links are treated as embeds
* ExtractorXML.java
    shouldExtract() - return true if content starts with "<?xml" only if it
    does not also contain "<!doctype html" or "<html" early in the content
2011-04-20 20:57:41 +00:00
nlevitt 6d6c1059a5 Use Recorder.getContentReplayCharSequence() instead of deprecated
getReplayCharSequence()
2011-04-20 19:38:12 +00:00
nlevitt 70d21ca2da HER-1820 more eager xml link extraction
* ExtractorXML.java
    instead of considering only strings that start with http(s):, consider all
    strings that match UriUtils.NAIVE_LIKELY_URI_PATTERN (and, as before,
    constitute the entirety of the xml tag content or attribute value)
* UriUtils.java
    - NAIVE_LIKELY_URI_PATTERN - add quotes to the excluded characters so it
      does not eat the closing quote when matching xml attribute values, and
      make visibility public for use in ExtractorXML
    - isLikelyUri() - refactor false positive check into new method
      isLikelyFalsePositive() so that it can be used in ExtractorXML avoiding
      redundant check against NAIVE_LIKELY_URI_PATTERN
2011-04-20 01:18:00 +00:00
gojomo 991a67e0e4 Post 3.1.0-beta
* **/pom.xml
    chnge version-identifier to "3.1.0-SNAPSHOT"
2011-04-15 22:28:52 +00:00
gojomo 52d00a9917 Prep for 3.1.0-beta
* **/pom.xml
    increment declared version
2011-04-15 18:09:43 +00:00
nlevitt 2df2584573 Fix HER-1880 robots.txt less specific Allow overrides more specific Disallow
* RobotsDirectives.java
    - use plain ConcurrentSkipListSet<String> instead of PrefixSet, to
      maintain complete list of prefixes, so we can know the longest match
    - allows(String) - return true if longest matching Allow prefix is longer
      than or equal to longest matching Disallow prefix
* RobotstxtTest.java
    flip expected result of test of generic Allow against specific Disallow
2011-04-14 02:25:43 +00:00
gojomo bd45720bae improved comments about new naming conventions/default-template 2011-04-11 22:49:01 +00:00
nlevitt f2ad45045b * FetchWhois.java
fix whois uri syntax - add missing "/" token
2011-04-07 17:17:48 +00:00
nlevitt 020ef44551 * FetchWhois.java
class-level javadoc including informal specification of whois uri
2011-04-07 06:44:06 +00:00
gojomo f4cf64af29 [HER-1053] compressed HTTP fetch: "Accept-encoding: gzip"
[HER-728] Offer replay stream that has been un-chunked (whether because response was HTTP/1.1 or used chunked in HTTP/1.0 against spec)
[HER-1876] Offer HTTP/1.1 option - for chunked transfer-encoding (but not persistent connections) 
* (all)
    update to use new (decoded-as-necessary) content streams/CharSequences
2011-04-06 20:54:16 +00:00
gojomo c5eb3ffd53 [HER-1053] compressed HTTP fetch: "Accept-encoding: gzip"
[HER-728] Offer replay stream that has been un-chunked (whether because response was HTTP/1.1 or used chunked in HTTP/1.0 against spec)
[HER-1876] Offer HTTP/1.1 option - for chunked transfer-encoding (but not persistent connections) 
* FetchHTTP
    add 'acceptCompression' and 'useHTTP11' properties, both default false
* Recorder
    track whether recorded-input is transfer-encoded (chunked) or content-encoded (gzip etc)
    offer alternate replay streams for 
      (1) raw 'messageBody'; 
      (2) entity (un-chunked if necessary)
      (3) content (decompressed if necessary)
    always content & GenericReplayCharSequence for CharSequence replays
* GenericReplayCharSequence
    always use a stream (rather than random-access buffer)
    always decode to prefix buffer first, so short content never touches disk no matter the encoding
* InMemoryReplayCharSequence
    deleted; 'Generic' now works similarly for small content and anyway random-access for single-byte-encodings is now rarely possible (given deconding streams)
* RecordingInputStream, RecordingOutputStream
    adjust for changed stream names, functionality moved to Recorder
* ReplayCharSequence
    use Charset instances rather than names
* ReplayInputStream
    add convenience constructor (and tmp-file-destroy) for copying any other inputStream into a seekable ReplayInputStream
2011-04-06 20:42:43 +00:00
nlevitt b79308b2e0 Part of HER-1875 - socket timeout for whois fetcher
* FetchWhois.java
2011-04-05 01:22:01 +00:00
nlevitt 81cda7be59 HER-1875 followup - override SocketFactory to support connect timeout for ftp
data connection
2011-04-05 01:21:09 +00:00
nlevitt 653f24e713 Avoid exception - "org.springframework.beans.NotWritablePropertyException:
Invalid property 'timeoutSeconds' of bean class
[org.archive.modules.fetcher.FetchFTP]: Bean property 'timeoutSeconds' is not
writable or has an invalid setter method. Does the parameter type of the setter
match the return type of the getter?"
* FetchFTP.java
    change setTimeoutSeconds(Integer) to setTimeoutSeconds(int)
2011-04-04 20:10:54 +00:00
nlevitt 8e7b9a5fdf Fix HER-1875 FetchFTP needs socket timeout
* FetchFTP.java
    setSoTimeoutMs(), getSoTimeoutMs()
    (init) - set default value of 20000 ms, same as FetchHTTP
    fetch() - set various timeouts on FTPClient - effectively sets connect
    timeout and socket timeout on control connection, and socket timeout on
    data connection (notably, does not set connect timeout on data connection
    w/ current commons-net)
2011-04-04 19:50:56 +00:00
nlevitt 013b5e394d Different fix for HER-1872: additional property preloadSourceUrl. Use either
that or preloadSource (a filesystem path), rather than try to handle both
possibilities with one string.
* PersistLoadProcessor.java
* PersistProcessor.java
2011-03-02 01:24:28 +00:00
nlevitt d97d4aa8cb Fix HER-1872 PersistLoadProcessor preloadSource does not support url
(regression from h1)
* PersistLoadProcessor.java
    change preloadSource from ConfigFile to String, allowing the logic in
    PersistProcessor.copyPersistSourceToHistoryMap() to decide how to load it
2011-02-26 00:59:35 +00:00
nlevitt b9c408cb07 Change form of generic serverless whois url to e.g. whois:archive.org
* FetchWhois.java
* UURIFactory.java
2011-02-19 02:00:18 +00:00
nlevitt 368ff65286 * FetchWhois.java
add license preamble
2011-02-17 22:52:37 +00:00
nlevitt 4991b295bc HER-1645 whois fetcher
* URIAuthorityBasedQueueAssignmentPolicy.java
    whois urls all go in special "whois" queue
* PreconditionEnforcer.java, CrawlURI.java
    move markPrerequisite() method from PreconditionEnforcer to CrawlURI
* profile-crawler-beans.cxml
    commented-out whois fetcher clause
* FetchStatusCodes.java
    new status codes for whois
* ServerCache.java
    getHostFor() - return "whois:" for whois urls, following dns: convention
* UURIFactory.java
    ugly hack for whois urls
* SchemeNotInSetDecideRule.java
    add whois to list of supported schemes
* WriterPoolProcessor.java
    shouldWrite() - we should write successful whois fetches
* WARCWriterProcessor.java
    write whois records
* FetchWhois.java
    the fetcher
2011-02-17 22:33:28 +00:00
gojomo 548114dbde [HER-1864] cookies are being set and sent back on TLDs/public-suffixes (like '.com') (sometimes called 'supercookies')
* CookieSpecBase
    use Guava InternetDomainName public-suffixes to both (1) prevent storing cookie on public-suffix; (2) prevent looking-up cookies for public-suffixes
* BdbCookieStorage
    correct checkpoint/relaunch behavior: don't reuse persisted cookies on normal relaunches; do reuse on resume-from-checkpoint
2011-02-16 01:18:57 +00:00
gojomo 40f7debeaa [HER-1846] IOException due to disk full situation can cause frontier lock-up
* WARCWriterProcessor
    move checkSize into block protected by catch/finally cleanup of errors
* WARCWriterProcessorTest
    test to simulate record-write-failure and verify does not cause pool lockup
* WARCWriterTest, WriterPoolMember
    comment and warnings cleanup
2011-01-28 03:02:17 +00:00
gojomo c1a99a8e35 [HER-1861] Add 'first-named' robots policy, which considers rules for more than one user-agent, in order
* Robotstxt, RobotstxtTest
    treat wildcardRules specially, so that it is possible to ask for only explicitly-named matches
* MostFavoredRobotsPolicy
    improve comment, change masquerade default to true
* FirstNamedRobotsPolicy, FirstNamedRobotsPolicyTest
    policy to consider alternate user-agents in order, using the first that has any explicit (non-wildcard) rules declared (or failing that, the original default user-agent)
2011-01-28 00:48:25 +00:00
gojomo a3d4dc5ac2 [HER-1834] log-bloat, memory problems when hopsPath grows endlessly with no max-hops limit set
* TooManyHopsDecideRule
    remove inadvertent reversal of sense of comparison, making default scope reject almost everything (and breaking most self-tests)
2011-01-26 03:08:32 +00:00
gojomo 8212d1c0dd [HER-1860] H3: Remove dependence on GC/finalization/PhantomReference magic in used ObjectIdentityCache implementation
* StatisticsTracker
    move small stat maps to in-memory ConcurrentMaps
    update processedSeedRecords for new ObjectIdentityCache shape
    move hostsDistribution/hostsBytes/hostsLastFinished tracking to CrawlHosts/hostsCache
* CrawlSummaryReport
    gets hosts count from hostsCache
* SeedRecord
    implement IdentityCacheable
* FetchStats
    remember last-success times
2011-01-20 23:36:50 +00:00
gojomo 0b577c1245 [HER-1860] H3: Remove dependence on GC/finalization/PhantomReference magic in used ObjectIdentityCache implementation
* IdentityCacheable
    new interface required of objects stored in ObjectIdentityCaches
* IdentityCacheableWrapper
    wrapper for storing arbitrary objects in ObjectIdentityCaches
* ObjectIdentityCache
    keys now always Strings
    values now always IdentityCacheables
    new dirtyKey() method to ensure a key is persisted
* ObjectIdentityBdbCache, ObjectIdentityMemCache
    update for new ObjectIdentityCache shape; still using legacy GC magic
* ObjectIdentityBdbManualCache
    alternate implementation relying on dirtying to ensure persistence
    use Guava library MapMaker for soft memMap; capped dirtyMap
* BdbModule
    use ObjectIdentityBdbManualCache by default
    up expectedConcurrency default to 64
* BdbModuleTest, ObjectIdentityBdbCacheTest, ObjectIdentityBdbManualCacheTest
    update, add tests
* Frontier
    FrontierGroup as IdentityCacheable
* AbstractFrontier, WorkQueueFrontier, BdbFrontier
    touchup queue/group accessors
    make queue instances dirty whenever mutated 
* WorkQueue, CrawlHost, CrawlServer
    IdentityCacheable support; appropriate makeDirty()s
* ServerCache, DefaultServerCache
    update for new ObjectIdentityCache shape
2011-01-20 23:28:38 +00:00
gojomo d2f28f051c [HER-1834] log-bloat, memory problems when hopsPath grows endlessly with no max-hops limit set
* CrawlURI
    add getHopCount() for total (including tallied overflow) hop count
* TooManyHopsDecideRule
    base decision on getHopCount()
2011-01-19 07:37:43 +00:00
gojomo 6795b3bfba [HER-1834] log-bloat, memory problems when hopsPath grows endlessly with no max-hops limit set
* CrawlURI
    MAX_HOPS_DISPLAYED constant, 50
    build new CrawlURI pathFromSeed value via extendHopsPath, which shows only MAX_HOPS_DISPLAYED hops, and precedes string with int value of hops unshown
* CrawlURITest
    test for extendHopsPath
2011-01-19 02:30:42 +00:00
gojomo 57edfde618 [HER-1859] ExtractorJS improper handling of Javascript string literal escapes
* ExtractorJS
    apply JS literal string deescaping before interpretation as link
* ExtractorJSTest
    3 basic tests, including one for this issue
* Link
    provide logging and fallback for case where via may be null; risk identified in creation of test cases could present again if a '.js' file is ever provided as a (via-less) seed
2011-01-18 23:34:35 +00:00
nlevitt 0af0d3cc6a * WARCWriterProcessor.java
addStats() - use putIfAbsent() to avoid race conditions; move
    creation of stats map to constructor to avoid another race condition
* WARCWriter.java
    switch back to thread-unsafe data structures for stats since object is used
    in one thread at a time according to javadoc
2011-01-18 04:02:42 +00:00
gojomo 823aaf0154 Add Javadoc class comment notes about (potentially) left-open-by-design ReplayCharSequence instances 2011-01-18 01:32:26 +00:00
nlevitt 69429f2d4e Missed one file r7051 commit
* ExtractorJS.java
    do not close ReplayCharSequence
2011-01-17 21:27:39 +00:00
nlevitt 50bb39a839 * WARCWriterProcessor.java
initialize urlsWritten (fix npe, oops)
2011-01-15 19:46:42 +00:00
nlevitt 72932b8ca8 Rename and comment on WARCWriter stats bits to clarify their use as temporary
accumulators on behalf of WARCWriterProcessor.
2011-01-15 04:17:49 +00:00