Commit Graph
406 Commits
Author SHA1 Message Date
Noah Levitt 4abfcb01b0 HER-1895 support HTTP Refresh header 2013-08-26 18:33:02 -07:00
Noah Levitt 68adaa70fd need to close the recorder here so it can calculate the content length properly 2013-08-09 19:14:58 -07:00
Noah Levitt 705a375daf uses of UriUtils.isLikelyUri() in Extractor{HTML,SWF,XML} with UriUtils.isVeryLikelyUri() to reap the benefits of HER-1523 improvements (should address archive-it issue ARI-3492) 2013-08-09 18:17:11 -07:00
Noah Levitt 8c64cba9ae TransclusionDecideRule.java - consider form submission "S" hop as navlink like "L", not transcluded 2013-08-05 16:05:02 -07:00
Noah Levitt f4f3b9be38 add logging to BdbUriUniqFilter.forgetSchemeHost(), clean up some javadocs 2013-07-26 17:50:36 -07:00
Noah Levitt 72094213a4 MirrorWriterProcessor.java - change "path" setting to a ConfigPath and set default value "${launchId}/mirror" 2013-07-22 17:54:06 -07:00
Adam Miller b85828135b Convert MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule to use CrawlURI.getFetchStatus() to avoid null status codes and broaden use.
Added test cases for MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule
Fix comments for youtube extractor.
2013-06-21 12:03:41 -07:00
Noah Levitt f2984881f6 address HER-2042 - new annotation duplicate:uriAgnosticDigest for all url agnostic duplicates, whether written to warcs or not; stats consult this annotation so that duplicate/novel numbers are correct 2013-06-07 17:17:24 -07:00
Noah Levitt ffd248f780 HER-2040 ExtractorJS shouldExtract when content-type is application/json 2013-06-03 11:23:16 -07:00
Noah Levitt d4c5bd6e98 avoid npe by setting default <input> type "text" when unspecified 2013-05-14 18:17:09 -07:00
Noah Levitt d972731958 ExtractorHTMLForms.java - handle a corner case extracting attribute values 2013-05-14 16:23:41 -07:00
Noah Levitt 7bd010e3c7 HTMLForm.java - default html <input> type is text, so consider <input> with no type as candidateUsernameInput 2013-05-13 22:53:37 -07:00
Noah Levitt bea5c6f267 followup on HER-2031 ExtractorHTMLForms.java - improve regexes looking for form and input attributes, mainly to avoid maxing cpu for long periods of time for certain input 2013-05-13 22:37:02 -07:00
Noah Levitt ec824e9724 avoid need for getRawData() method by introducing Link.hasData() 2013-05-13 17:43:45 -07:00
Kristinn Sigurðsson 7c5b2ceb00 Added a data map to Link. Becomes the CrawlURI's data map when Link is
promoted to CrawlURI
2013-05-13 15:18:02 +00:00
Noah Levitt d5693563f8 fix bug pointed out in https://webarchive.jira.com/browse/ARI-3176?focusedCommentId=35314 - url-agnostic dedupe not reported in host report (and other reports) 2013-04-29 18:11:24 -07:00
Noah Levitt 6a0ed067c0 consolidate disparate versions of WARCConstants into one version that lives in ia-web-commons 2013-03-29 17:12:32 -07:00
Kenji Nagahashi a4a3ce60cd fix failing ContentDigestHistoryTest 2013-03-26 15:06:31 -07:00
Noah Levitt 5d8fe0443d Merge branch 'master' into HER-2031 2013-03-23 19:16:05 -07:00
Noah Levitt 8fd361e518 update source/target version to 1.6 since we use 1.6 features, specifically @Override for methods only implementing an interface (not sure why maven wasn't enforcing this); also remove maven-antrun-plugin created timestamp.txt, doesn't seem to be used anywhere 2013-03-20 14:25:22 -07:00
Noah Levitt c509b6e552 remove more eclipse stuff, people like to manage their own 2013-03-20 13:39:28 -07:00
Noah Levitt b0c6619f80 Merge branch 'master' of github.com:internetarchive/heritrix3 2013-03-18 20:02:03 -07:00
Noah Levitt 0c9e2a1f4d ContentDigestHistory{Loader,Storer} - ignore uris with empty content body (deduping that is counterproductive) 2013-03-18 20:01:52 -07:00
Noah Levitt ff36546eda fix serverless whois url example in comment 2013-03-18 14:03:39 -07:00
gojomo c5c54ac138 comment-out debug output 2013-03-12 11:42:47 -07:00
gojomo 7c8295ee82 checkpoint support 2013-03-12 11:42:29 -07:00
gojomo a844878e7d Merge branch 'simple_logins' into HER-2031 2013-03-12 01:16:31 -07:00
gojomo 3f8590fe44 when found in CrawlURI, add WARC-Simple-Form-Province-Status header 2013-03-11 20:57:41 -07:00
gojomo bf9cab8eb2 when POST/submit-data requested, alter request method accordingly 2013-03-11 20:57:09 -07:00
gojomo bbac7d7320 per (overlay) settings, spawn high-priority POST when appropriate; annotate/add-headers to describe 2013-03-11 20:56:31 -07:00
gojomo ed639c8342 extract HTMLForm INPUT details 2013-03-11 20:55:20 -07:00
gojomo 94a61b4fb7 retain offsets of discovered FORMs to aid later processing 2013-03-11 20:53:41 -07:00
gojomo 6b0abae8ee new constants, hop-type; refactoring to help support form-submissions 2013-03-11 20:52:59 -07:00
Noah Levitt 8cbb84b8df Merge branch 'master' of github.com:internetarchive/heritrix3 2013-02-22 15:43:30 -08:00
Noah Levitt fb758e53e3 Avoid pulling in hadoop-core.jar and its dependencies; add plugin versions to get rid of warnings 2013-02-22 15:43:15 -08:00
Noah Levitt 878e3b1afe Remove spurious "throws" declaration 2013-02-05 18:37:22 -08:00
Noah Levitt 543e8820ef Merge branch 'master' into archive-commons-refactor 2012-12-20 02:44:57 -08:00
Noah Levitt 3a65fa25d7 Working on refactoring some stuff into archive-commons 2012-12-18 14:46:52 -08:00
Noah Levitt 0eac126f7d Special WARC-Profile for uri-agnostic revisit records. See HER-2022, http://tech.groups.yahoo.com/group/archive-crawler/message/7806 2012-12-13 14:11:14 -08:00
Noah Levitt 5586d5452d * JerichoExtractorHTMLTest.java
avoid redundancy by using extractor built in setUp()
    makeExtractor() - call extractor.setExtractorJS(new ExtractorJS()) since we got rid of the static ExtractorJS stuff
    testConditionalComment1() - override to skip the test since it fails with JerichoExtractorHTML
2012-12-11 17:00:23 -08:00
Noah Levitt 85118c59c4 Tighten up javascript extraction based on real-world analysis (addresses HER-1523)
* UriUtils.java
    new method isVeryLikelyUri() with tighter heuristic than isLikelyUri()
* ExtractorJS.java
    use UriUtils.isVeryLikelyUri(), and change order of operations to do fixup before call to isVeryLikelyUri(), since it doesn't expect strings with javascript escaping and stuff
* StringExtractorTestBase.java
    handle test data with expected value null, meaning no outlinks expected
* ExtractorHTMLTest.java
    avoid redundancy by using the extractor created in ContentExtractorTestBase.setUp()
* ExtractorJSTest.java
    some new tests
2012-12-11 16:27:26 -08:00
Noah Levitt b1cd3d30a0 * ExtractorHTMLTest.java
makeExtractor() - call setExtractorJS(new ExtractorJS()) so testSpeculativeLinkExtraction() passes
2012-12-10 17:31:55 -08:00
Noah Levitt 256b16249a * ExtractorJS.java
refactor so considerStrings() is not static, allowing it to be overridden in subclasses
* ExtractorHTML.java, ExtractorJS.java
    add @Autowired parameter extractorJS used to process inline javascript, instead of call to static ExtractorJS.considerStrings()
2012-12-10 16:51:42 -08:00
Noah Levitt e44cf9ca9b A little cleanup of FetchHTTPTest refactoring 2012-12-04 18:43:15 -08:00
Noah Levitt 5354960f00 Refactor FetchHTTPTest to use a TestSuite so that starting and shutting down the test http servers happen only once at start and finish respectively 2012-12-04 17:25:09 -08:00
Noah Levitt e30cb83540 * FetchHTTPTest.java
testLaxUrlEncoding() - Tests a URL not correctly url-encoded, but that heritrix lets pass through to mimic browser behavior.
2012-12-04 16:25:09 -08:00
Noah Levitt 92b302f164 * ExtractorMultipleRegexTest.java
avoid "constant string too long" compile error
2012-11-16 19:08:18 -08:00
Noah Levitt 94df568b4f * ExtractorMultipleRegex.java
remove temporary performance testing code
2012-11-16 18:34:28 -08:00
Noah Levitt aa98ec3327 * ExtractorMultipleRegex.java
keep cache of groovy Template objects, since they are expensive to create
2012-11-16 18:30:41 -08:00
Noah Levitt eeee01a52d * ExtractorMultipleRegex.java
javadocs
2012-11-16 16:42:43 -08:00