Commit Graph
719 Commits
Author SHA1 Message Date
nlevitt e6de9cd96b * Sheet.java
prime() - fix exception triggered in process of handling earlier real exception, which was masking the real exception - see http://tech.groups.yahoo.com/group/archive-crawler/message/7303
2011-08-30 22:41:37 +00:00
nlevitt cdb6fb3a9a * PersistProcessor.java
copyPersistEnv() - improve logging on history entries by marshalling them as json
2011-08-30 22:04:17 +00:00
nlevitt c70f7ccbc0 * profile-crawler-beans.cxml
remove reference to defunct frontier setting "holdQueues", pointed out by Kristinn Sigurðsson
2011-08-08 23:42:47 +00:00
nlevitt 28aa2ffede * **/pom.xml
improve formatting (no substantive changes at all -- try "svn diff -x -ub -c7232")
2011-08-05 17:46:57 +00:00
nlevitt 178fdd92b6 * **/pom.xml
remove references to unused, defunct agilejava maven repository (had been used for xdoc stuff)
2011-08-05 17:37:16 +00:00
nlevitt ada8cbd198 Fix bug reported by Travis where in certain corner cases where urls are processed in a particular order, a seed can be incorrectly tagged with S_ROBOTS_PREREQUISITE_FAILURE
* DispositionProcessor.java
    innerProcess() - do not update server's robots info if robots fetch has been deferred
2011-08-04 01:32:09 +00:00
nlevitt 706869e586 Fix handling of server error on robots.txt - change to what appears to be the intended behavior
* DispositionProcessor.java
    innerProcess() - call server.updateRobots(curi) after special ignore-robots short-circuiting of retry cycle rather than before, so that updateRobots() deemed-not-found handling applies
* FetchHTTP.java
    call curi.setHttpMethod() right after creating the HttpMethod, instead of when processing the response; this is needed so that updateRobots() check of fetchType works properly (before this change, fetchType was "UNKNOWN" on server error)
* PreconditionEnforcer.java, CrawlServer.java
    fix typos
2011-08-02 23:28:20 +00:00
nlevitt 3cfc315de9 Fix bug reported by Travis where each url in action directory ".include" file is reported as "not canonicalized" in heritrix_out, and some end up crawled in spite of being considered included
* WorkQueueFrontier.java
    considerIncluded(CrawlURI) - call FrontierPreparer.prepare(CrawlURI) so that CrawlURI is canonicalized etc before being considered included (similar to what schedule() does)
* CrawlURI.java
    getCanonicalString(), getPolitenessDelay() - use logger instead of System.err
2011-08-02 01:07:11 +00:00
nlevitt dcbc259776 Fix HER-1917 failure to use good http authentication type if server supports any that heritrix doesn't - problem reported and fix suggested by Adam Wilmer
* HttpAuthenticationCredential.java
    populate() 
        +  http.getParams().setParameter(AuthPolicy.AUTH_SCHEME_PRIORITY,
        +          Arrays.asList(AuthPolicy.DIGEST, AuthPolicy.BASIC));
2011-07-30 01:30:44 +00:00
nlevitt 26584d44d7 * ExtractorHTMLTest.java
testScriptTagWritingScriptType() - As of r7225, ExtractorJS finds a new url - avoid using a strategically placed space. (The fact that UriUtils.isLikelyUri() returns true for the string ");document.write(unescape(" is kind of crazy, but a separate issue)
2011-07-27 20:04:54 +00:00
nlevitt bda2fbebe4 Fix HER-1873 strings with spaces can confuse ExtractorJS
* ExtractorJS.java
    considerStrings() - backtrack to just ahead of previous closing quote to reconsider it as a possible opening quote
* ExtractorJSTest.java
    minimal test case
2011-07-27 18:56:37 +00:00
nlevitt b2307e7cf1 Fix HER-1914 cookie loading and saving does not work properly
* AbstractCookieStorage.java
    make cookiesLoadFile a ConfigFile and use obtainReader() so that the file is snapshotted into launch dir
    saveCookies() - include expiration date, so file format match the format it claims to use, and what loadCookies() expects 
    loadCookies() - handle expiration date field, and actually put each cookie in the cookie map
    refactor loadCookies() for use with Reader from ConfigFile.obtainReader()
    fix up javadocs some
2011-07-26 19:30:16 +00:00
nlevitt 62e1c835dd Fix HER-1913 cookies not being sent
* CookieSpecBase.java
    match(String, int, String, boolean, SortedMap) - use InternetDomainName.name() instead of .toString(), since the latter doesn't return the plain old domain name
* Cookie.java
    javadoc typo
2011-07-26 19:07:26 +00:00
nlevitt cb7a454d0c Post 3.1.0-RC1
* **/pom.xml
   switch version back to "3.1.0-SNAPSHOT"
2011-07-26 03:42:02 +00:00
nlevitt 347a25e417 Prep for 3.1.0-RC1 release
* pom.xml, */pom.xml
    bump version number to 3.1.0-RC1
2011-07-26 02:57:48 +00:00
nlevitt 3a4419967d Remove PathFixupListener, which was only used for SurtPrefixedDecideRule initialization, now done differently as of r7218.
* ConfigPathConfigurer.java
    remove calls to PathFixupListener
* PathFixupListener.java
    removed
2011-07-25 23:10:02 +00:00
nlevitt a103c5f05c * SurtPrefixedDecideRule.java
fix timing reading surts source to make sure launch dir is in place for snapshot
2011-07-25 22:52:15 +00:00
nlevitt 0c8840db56 * .classpath
reference more source jars retrieved with "mvn dependency:sources"
2011-07-25 21:53:50 +00:00
nlevitt efff6feed6 * CrawlJob.java
avoid NPE on build-then-terminate without a launch
2011-07-25 19:16:31 +00:00
nlevitt e2f29992d4 Put hard limit of 5 on seed redirects being treated as new seeds, to avoid
endless recrawling of the same url(s) in case of a redirect loop.  
* CandidatesProcessor.java
2011-07-23 02:20:35 +00:00
gojomo 269a763860 [HER-1912] Kryo-deserialization of Robotstxt can consume excess memory by creating multiple copies of RobotsDirective
* CrawlServer, Robotstxt
    move autoregister details to Robotstxt, RobotsDirectives
* RobotsDirectives
    use ReferenceFieldSerializer so that repeated instances of RobotsDirectives in same context are replaced with backrefs
* RobotstxtTest
    (from Kenji) unit test for instance multiplication
2011-07-22 06:46:39 +00:00
gojomo a490737b59 * StatisticsTracker
add 'trackSources' to make hosts-by-source-tag report optional in large crawls, where it is very expensive (and an OOME risk until other changes are also made)
2011-07-20 23:36:19 +00:00
gojomo a69126ad21 Reduce redundant Pattern instance caching
* TextUtils
    maintain global Pattern soft-cache by regex string key
2011-07-20 23:04:26 +00:00
nlevitt b3e35499ca HER-1812 Spring-ify how reports are handled - based partly on patch from Kristinn Sigurðsson
* Report.java
    make StatisticsTracker stats an argument to write() instead of a field; new properties shouldReportAtEndOfCrawl and shouldReportDuringCrawl, both default true, but configurable if desired, replacing the LIVE_REPORTS/END_REPORTS stuff in StatisticsTracker
* StatisticsTracker.java
    new property "reports", a List<Report>, which defaults (with lazy initialization) to the same reports we've always had, replacing hard-coded lists of reports classes and related code
* profile-crawler-beans.cxml
    default reports list, commented out
* SeedsReport.java FrontierSummaryReport.java MimetypesReport.java SourceTagsReport.java FrontierNonemptyReport.java CrawlSummaryReport.java ResponseCodeReport.java HostsReport.java ProcessorsReport.java ToeThreadsReport.java
    update write(PrintWriter, StatisticsTracker) arguments
* JobResource.java
    replace use of StatisticsTracker.LIVE_REPORTS
2011-07-20 17:05:12 +00:00
nlevitt e9d53ec487 HER-1791 - a different implementation
* IdenticalDigestDecideRule.java
    fix javadoc typo
* CoreAttributeConstants.java
    "history-good-to-store" is history
* RecrawlAttributeConstants.java
    new "write-tag" constant
* PersistProcessor.java
    remove defunct "history-good-to-store" logic; nadd new option onlyStoreIfWriteTagPresent, and update shouldStore() to respect that option
* WriterPoolProcessor.java
    remove defunct "history-good-to-store" logic; new method copyForwardWriteTagIfDupe(); override innerRejectProcess() to call copyForwardWriteTagIfDupe()
* ARCWriterProcessor.java, WARCWriterProcessor.java
    remove defunct "history-good-to-store" logic; add writeTag to history on write; call copyForwardWriteTagIfDupe() on skipIdenticalDigest
2011-07-18 20:22:48 +00:00
gojomo 421164aaab * ConfigPathConfigurer
log as WARNING when snapshot not possible because launch directory not yet available
2011-07-16 17:49:11 +00:00
nlevitt 0676ded3f2 * CrawlJob.java
fix logging of job events
2011-07-16 00:48:14 +00:00
nlevitt f13b393dab * PersistLogProcessor.java
put default persist log path in launch directory
2011-07-16 00:44:13 +00:00
gojomo e819d197fe [HER-915] (contrib) add http proxy authentication
contributed by Adam Wilmer
* FetchHTTP.java
    properties, setup for http proxy user/password
* profile-crawler-beans.cxml
    commented-out example settings for proxy auth properties
2011-07-15 21:57:20 +00:00
gojomo 4221aee537 fix selftests
* selftest-crawler-beans.cxml
    write logs, arcs to old constant locations, for ease of post-test verification
2011-07-15 21:34:59 +00:00
nlevitt b90aa89e4b More on HER-1901 - fix build by refactoring launch dir initialization into PathSharingContext; refactor config path interpolation mostly into ConfigPathConfigurer; various tweaks
* ActionDirectory.java, CrawlerLoggerModule.java, StatisticsTracker.java, SurtPrefixedDecideRule.java, WriterPoolProcessor.java, profile-crawler-beans.cxml
    use camelcase ${launchId} 
* CheckpointService.java, JobResource.java
    rename getAvailableCheckpointDirectories() to findAvailableCheckpointDirectories() so that it's not treated as a bean property
* WriterPoolProcessor.java, WriterPoolSettingsData.java, WriterPoolMember.java
    rename getOutputDirs() to calcOutputDirs() so that it's not treated as a bean property
* ConfigPath.java
    let ConfigPathConfigurer to do interpolation of ${launchId}
* ConfigFile.java
    let ConfigPathConfigurer do the snapshotting of config file
* CrawlJob.java, PathSharingContext.java
    remove launchDir initialization out of CrawlJob into PathSharingContext
* ConfigPathConfigurer.java
    - remove special handling of WriterPoolProcessor store paths; instead, look in beans for ConfigPaths within Iterables
    - have each ConfigPath hold reference to this ConfigPathConfigurer to use for interpolating ${launchId} and snapshotting config files
2011-07-15 19:39:32 +00:00
gojomo 1bfbec8772 fix build; move Preformatter to package (commons) where first referenced 2011-07-15 19:02:14 +00:00
gojomo 2e10d07da2 [HER-741] Make extractors interrogate for charset
* ExtractorHTML
    fix reflexive probe to look as many characters in (1000) as initially
2011-07-14 06:53:56 +00:00
gojomo be12d1e8f3 calmer logging (fewer alerts in job.log/alerts.log/heritrix_out.log)
* Recorder
    object to unsupported content-encodings when first set
* FetchHTTP
    note as annotation unsupported content-encodings
* Link
    not as annotation when base URI is used for absent via
* ExtractorHTML
    downgrade char-sequence reading problem (often a chunking problem) to WARNING from SEVERE
2011-07-14 06:35:07 +00:00
gojomo 5e2e458469 * StoredQueue
(size) more robust against concurrent emptying
2011-07-14 06:24:14 +00:00
nlevitt 9eb0c137e8 HER-1901 twiddles
* CrawlJob.java
    decided to go with plain timestamp17 as launch id, no "launch-" prefix
* ConfigPathConfigurer.java
    refactored special remembering of WriterPoolProcessor storePaths out of fixupPaths()
2011-07-14 02:03:01 +00:00
gojomo 6a3072ef46 more rapidly release refs for finished CrawlURIs
* UriProcessingFormatter, Preformatter, GenerationFileHandler
    improve preformat-outside-synchronized optimization so that the LogRecord/CrawlURI doesn't linger until next displaces it
* CrawlURI
    (processingCleanup) null more of last-processing-run collected values
2011-07-14 01:14:04 +00:00
nlevitt c84e87c648 More HER-1901
* CrawlJob.java 
    job log in launch dir covering just that launch
2011-07-13 20:47:08 +00:00
nlevitt 4d966ae70b More on HER-1901 - support ${launch-id} interpolation on W/ARCWriterProcessor storePaths
* WriterPoolProcessor.java, ARCWriterProcessor.java, WARCWriterProcessor.java
    change type of storePaths to List<ConfigPath> and handle appropriately
* ConfigPathConfigurer.java
    fixupPaths() - old code did not touch WriterPoolProcessor storePaths, since they're deeply nested inside the bean, but they need to be remembered for later interpolation of ${launch-id}, so add special handling
2011-07-13 20:18:51 +00:00
nlevitt d5b299d1ef * CrawlJob.java
initLaunchId() - didn't mean for this method to be public
2011-07-13 20:13:13 +00:00
nlevitt c39c927447 * profile-crawler-beans.cxml
runWhileEmpty defaults to false
2011-07-13 19:20:45 +00:00
nlevitt 4250e07de9 HER-1901 timestamped subdirectory for each launch
* HardLinker.java
    renamed FilesystemLinkMaker.java
* FilesystemLinkMaker.java
    add support for symbolic links
* CLibrary.java
    new method symlink()
* BdbModule.java
    use new class name FilesystemLinkMaker
* CrawlJob.java
    at crawl launch, create launch directory launch-{timestamp17}, copy cxml there, symlink "current" to launch dir, inform ConfigPaths
* ConfigPath.java
    interpolate ${launch-id} in configured paths
* ConfigFile.java
    obtainReader() - snapshot config files to launch dir when they are read
* ActionDirectory.java
    default doneDir now ${launch-id}/actions-done
    actOn() - symlink from old style done dir action/done to done files
* SurtPrefixedDecideRule.java
    default surtsDumpFile now ${launch-id}/surts.dump
    pathsFixedUp() - this gets called at build time, but we don't want anything written to disk until launch time, so remove call to dumpSurtPrefixSet() here
* CrawlerLoggerModule.java
    default logs dir now ${launch-id}/logs
* StatisticsTracker.java 
    default reports dir now ${launch-id}/reports
* WriterPoolProcessor.java
    default writer base path now ${launch-id} 
* profile-crawler-beans.cxml
    update to reflect new default paths under launch dirs
* PropertyUtils.java
    fix javadoc typo
2011-07-13 19:18:55 +00:00
gojomo a8682ff1ff * TextSeedModule
(actOn) use ArchiveUtils.getBufferedReader so that seeds-files via action directory may optionally be gzipped (indicated by .gz suffix)
2011-07-09 02:13:52 +00:00
gojomo d742434b4e [HER-1911] make writer-pool 'maxActive' setting adjustable mid-crawl
* WriterPool
    make new LARGEST_MAX_ACTIVE (255) the capacity of the reuse-pool, so that the maxActive property may vary up to that and still work; also handle more gracefully the failure of trying to use a larger number than the capacity (writer gets closed early rather than left hanging open)
2011-07-08 21:59:05 +00:00
gojomo a41809b8a7 [HER-1908] H3: upgrade to Spring 3
* pom.xml, .classpath
    update references to necessary spring-3.0.5 packages
* SeedModule
    merge rather than replace event listeners, so that (now later) autowiring doesn't clobber the self-insertion done by anonymous (non-top-level) SurtPrefixedDecideRule bean
* HeritrixLifecycleProcessor, PathScharingContext
    use this new non-default LifecycleProcessor to avoid new Spring3 behavior of auto-start()ing context on refresh(build)
* CrawlController
    example of using @Value annotation to set default value: will offer benefits for auto-discovery of defaults for configuration interface, or enforcing maximally-explicit configurations (see [HER-1897])
* profile-crawler-beans.cxml, selftest-crawler-beans.cxml
    update templates with spring3 preamble/boilerplate
2011-07-08 04:31:54 +00:00
gojomo aa15d293e8 [HER-1909] H3: slow checkpoint resume, locking out web UI during resume
[HER-1910] H3: job-related pages very slow to render (pause mid-way through) in large crawls
* BdbModule
    option to reuse bdb data on StoredQueue-creation
* Checkpoint
    new saveWriter, loadReader utilities for creating extra checkpoint files (for the long list of ready queues)
* BdbFrontier
    on checkpoint, remember ordered list of inProcess/ready/snoozed queues (so same queues are active upon resume)
    on resume, reuse old retired/inactive StoredQueue data, and reload active queues from newly-saved list
* StoredQueue
    restore tailIndex properly when resuing old data
    more-efficient size() calculation (for HER-1910)
2011-07-08 04:30:32 +00:00
gojomo 2b0d967b32 Resolve 'Browse Beans' issue after built, pre-launch
* JobRelatedResource
    catch InvalidPropertyException, render as red error message rather than fouling entire trace of beans to web UI
* StatisticsTracker
    rename getSnapshots to listSnapshots to prevent interpretation as (invalid) bean-property
2011-06-27 22:25:12 +00:00
nlevitt 041554fa5d Fix for HER-1906 checkpointing gives error on Windows
* HardLinker.java
    makeHardLink() - wrap unix link() and windows CreateHardLink() and choose depending on platform
* BdbModule.java
    doCheckpoint() - use HardLinker.makeHardLink()
2011-06-24 19:24:28 +00:00
gojomo 39091dc82c Fix tests
* TextSeedModule
    remove unintended @Required annotation
2011-06-22 23:09:43 +00:00
gojomo 842f495ee9 [HER-1905] H3: allow crawling to begin while large seed list still loading
* TextSeedModule
    refactor announceSeeds to occur in background thread if non-default blockAwaitingSeedLines value is set
    signal CountDownLatch on each line, allowing calling thread to proceed at right count
* profile-crawler-beans.cxml
    commented-out blockAwaitingSeedLines default settings (-1, meaning wait for all seed lines)
2011-06-22 21:21:26 +00:00