Noah Levitt
4abfcb01b0
HER-1895 support HTTP Refresh header
2013-08-26 18:33:02 -07:00
Noah Levitt
68adaa70fd
need to close the recorder here so it can calculate the content length properly
2013-08-09 19:14:58 -07:00
Noah Levitt
705a375daf
uses of UriUtils.isLikelyUri() in Extractor{HTML,SWF,XML} with UriUtils.isVeryLikelyUri() to reap the benefits of HER-1523 improvements (should address archive-it issue ARI-3492)
2013-08-09 18:17:11 -07:00
Noah Levitt
8c64cba9ae
TransclusionDecideRule.java - consider form submission "S" hop as navlink like "L", not transcluded
2013-08-05 16:05:02 -07:00
Noah Levitt
f4f3b9be38
add logging to BdbUriUniqFilter.forgetSchemeHost(), clean up some javadocs
2013-07-26 17:50:36 -07:00
Noah Levitt
72094213a4
MirrorWriterProcessor.java - change "path" setting to a ConfigPath and set default value "${launchId}/mirror"
2013-07-22 17:54:06 -07:00
Adam Miller
b85828135b
Convert MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule to use CrawlURI.getFetchStatus() to avoid null status codes and broaden use.
...
Added test cases for MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule
Fix comments for youtube extractor.
2013-06-21 12:03:41 -07:00
Noah Levitt
f2984881f6
address HER-2042 - new annotation duplicate:uriAgnosticDigest for all url agnostic duplicates, whether written to warcs or not; stats consult this annotation so that duplicate/novel numbers are correct
2013-06-07 17:17:24 -07:00
Noah Levitt
ffd248f780
HER-2040 ExtractorJS shouldExtract when content-type is application/json
2013-06-03 11:23:16 -07:00
Noah Levitt
d4c5bd6e98
avoid npe by setting default <input> type "text" when unspecified
2013-05-14 18:17:09 -07:00
Noah Levitt
d972731958
ExtractorHTMLForms.java - handle a corner case extracting attribute values
2013-05-14 16:23:41 -07:00
Noah Levitt
7bd010e3c7
HTMLForm.java - default html <input> type is text, so consider <input> with no type as candidateUsernameInput
2013-05-13 22:53:37 -07:00
Noah Levitt
bea5c6f267
followup on HER-2031 ExtractorHTMLForms.java - improve regexes looking for form and input attributes, mainly to avoid maxing cpu for long periods of time for certain input
2013-05-13 22:37:02 -07:00
Noah Levitt
ec824e9724
avoid need for getRawData() method by introducing Link.hasData()
2013-05-13 17:43:45 -07:00
Kristinn Sigurðsson
7c5b2ceb00
Added a data map to Link. Becomes the CrawlURI's data map when Link is
...
promoted to CrawlURI
2013-05-13 15:18:02 +00:00
Noah Levitt
d5693563f8
fix bug pointed out in https://webarchive.jira.com/browse/ARI-3176?focusedCommentId=35314 - url-agnostic dedupe not reported in host report (and other reports)
2013-04-29 18:11:24 -07:00
Noah Levitt
6a0ed067c0
consolidate disparate versions of WARCConstants into one version that lives in ia-web-commons
2013-03-29 17:12:32 -07:00
Kenji Nagahashi
a4a3ce60cd
fix failing ContentDigestHistoryTest
2013-03-26 15:06:31 -07:00
Noah Levitt
5d8fe0443d
Merge branch 'master' into HER-2031
2013-03-23 19:16:05 -07:00
Noah Levitt
8fd361e518
update source/target version to 1.6 since we use 1.6 features, specifically @Override for methods only implementing an interface (not sure why maven wasn't enforcing this); also remove maven-antrun-plugin created timestamp.txt, doesn't seem to be used anywhere
2013-03-20 14:25:22 -07:00
Noah Levitt
c509b6e552
remove more eclipse stuff, people like to manage their own
2013-03-20 13:39:28 -07:00
Noah Levitt
b0c6619f80
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-03-18 20:02:03 -07:00
Noah Levitt
0c9e2a1f4d
ContentDigestHistory{Loader,Storer} - ignore uris with empty content body (deduping that is counterproductive)
2013-03-18 20:01:52 -07:00
Noah Levitt
ff36546eda
fix serverless whois url example in comment
2013-03-18 14:03:39 -07:00
gojomo
c5c54ac138
comment-out debug output
2013-03-12 11:42:47 -07:00
gojomo
7c8295ee82
checkpoint support
2013-03-12 11:42:29 -07:00
gojomo
a844878e7d
Merge branch 'simple_logins' into HER-2031
2013-03-12 01:16:31 -07:00
gojomo
3f8590fe44
when found in CrawlURI, add WARC-Simple-Form-Province-Status header
2013-03-11 20:57:41 -07:00
gojomo
bf9cab8eb2
when POST/submit-data requested, alter request method accordingly
2013-03-11 20:57:09 -07:00
gojomo
bbac7d7320
per (overlay) settings, spawn high-priority POST when appropriate; annotate/add-headers to describe
2013-03-11 20:56:31 -07:00
gojomo
ed639c8342
extract HTMLForm INPUT details
2013-03-11 20:55:20 -07:00
gojomo
94a61b4fb7
retain offsets of discovered FORMs to aid later processing
2013-03-11 20:53:41 -07:00
gojomo
6b0abae8ee
new constants, hop-type; refactoring to help support form-submissions
2013-03-11 20:52:59 -07:00
Noah Levitt
8cbb84b8df
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-02-22 15:43:30 -08:00
Noah Levitt
fb758e53e3
Avoid pulling in hadoop-core.jar and its dependencies; add plugin versions to get rid of warnings
2013-02-22 15:43:15 -08:00
Noah Levitt
878e3b1afe
Remove spurious "throws" declaration
2013-02-05 18:37:22 -08:00
Noah Levitt
543e8820ef
Merge branch 'master' into archive-commons-refactor
2012-12-20 02:44:57 -08:00
Noah Levitt
3a65fa25d7
Working on refactoring some stuff into archive-commons
2012-12-18 14:46:52 -08:00
Noah Levitt
0eac126f7d
Special WARC-Profile for uri-agnostic revisit records. See HER-2022, http://tech.groups.yahoo.com/group/archive-crawler/message/7806
2012-12-13 14:11:14 -08:00
Noah Levitt
5586d5452d
* JerichoExtractorHTMLTest.java
...
avoid redundancy by using extractor built in setUp()
makeExtractor() - call extractor.setExtractorJS(new ExtractorJS()) since we got rid of the static ExtractorJS stuff
testConditionalComment1() - override to skip the test since it fails with JerichoExtractorHTML
2012-12-11 17:00:23 -08:00
Noah Levitt
85118c59c4
Tighten up javascript extraction based on real-world analysis (addresses HER-1523)
...
* UriUtils.java
new method isVeryLikelyUri() with tighter heuristic than isLikelyUri()
* ExtractorJS.java
use UriUtils.isVeryLikelyUri(), and change order of operations to do fixup before call to isVeryLikelyUri(), since it doesn't expect strings with javascript escaping and stuff
* StringExtractorTestBase.java
handle test data with expected value null, meaning no outlinks expected
* ExtractorHTMLTest.java
avoid redundancy by using the extractor created in ContentExtractorTestBase.setUp()
* ExtractorJSTest.java
some new tests
2012-12-11 16:27:26 -08:00
Noah Levitt
b1cd3d30a0
* ExtractorHTMLTest.java
...
makeExtractor() - call setExtractorJS(new ExtractorJS()) so testSpeculativeLinkExtraction() passes
2012-12-10 17:31:55 -08:00
Noah Levitt
256b16249a
* ExtractorJS.java
...
refactor so considerStrings() is not static, allowing it to be overridden in subclasses
* ExtractorHTML.java, ExtractorJS.java
add @Autowired parameter extractorJS used to process inline javascript, instead of call to static ExtractorJS.considerStrings()
2012-12-10 16:51:42 -08:00
Noah Levitt
e44cf9ca9b
A little cleanup of FetchHTTPTest refactoring
2012-12-04 18:43:15 -08:00
Noah Levitt
5354960f00
Refactor FetchHTTPTest to use a TestSuite so that starting and shutting down the test http servers happen only once at start and finish respectively
2012-12-04 17:25:09 -08:00
Noah Levitt
e30cb83540
* FetchHTTPTest.java
...
testLaxUrlEncoding() - Tests a URL not correctly url-encoded, but that heritrix lets pass through to mimic browser behavior.
2012-12-04 16:25:09 -08:00
Noah Levitt
92b302f164
* ExtractorMultipleRegexTest.java
...
avoid "constant string too long" compile error
2012-11-16 19:08:18 -08:00
Noah Levitt
94df568b4f
* ExtractorMultipleRegex.java
...
remove temporary performance testing code
2012-11-16 18:34:28 -08:00
Noah Levitt
aa98ec3327
* ExtractorMultipleRegex.java
...
keep cache of groovy Template objects, since they are expensive to create
2012-11-16 18:30:41 -08:00
Noah Levitt
eeee01a52d
* ExtractorMultipleRegex.java
...
javadocs
2012-11-16 16:42:43 -08:00