Kristinn Sigurðsson
7c5b2ceb00
Added a data map to Link. Becomes the CrawlURI's data map when Link is
...
promoted to CrawlURI
2013-05-13 15:18:02 +00:00
Noah Levitt
d5693563f8
fix bug pointed out in https://webarchive.jira.com/browse/ARI-3176?focusedCommentId=35314 - url-agnostic dedupe not reported in host report (and other reports)
2013-04-29 18:11:24 -07:00
Noah Levitt
6a0ed067c0
consolidate disparate versions of WARCConstants into one version that lives in ia-web-commons
2013-03-29 17:12:32 -07:00
Kenji Nagahashi
a4a3ce60cd
fix failing ContentDigestHistoryTest
2013-03-26 15:06:31 -07:00
Noah Levitt
5d8fe0443d
Merge branch 'master' into HER-2031
2013-03-23 19:16:05 -07:00
Noah Levitt
8fd361e518
update source/target version to 1.6 since we use 1.6 features, specifically @Override for methods only implementing an interface (not sure why maven wasn't enforcing this); also remove maven-antrun-plugin created timestamp.txt, doesn't seem to be used anywhere
2013-03-20 14:25:22 -07:00
Noah Levitt
c509b6e552
remove more eclipse stuff, people like to manage their own
2013-03-20 13:39:28 -07:00
Noah Levitt
b0c6619f80
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-03-18 20:02:03 -07:00
Noah Levitt
0c9e2a1f4d
ContentDigestHistory{Loader,Storer} - ignore uris with empty content body (deduping that is counterproductive)
2013-03-18 20:01:52 -07:00
Noah Levitt
ff36546eda
fix serverless whois url example in comment
2013-03-18 14:03:39 -07:00
gojomo
c5c54ac138
comment-out debug output
2013-03-12 11:42:47 -07:00
gojomo
7c8295ee82
checkpoint support
2013-03-12 11:42:29 -07:00
gojomo
a844878e7d
Merge branch 'simple_logins' into HER-2031
2013-03-12 01:16:31 -07:00
gojomo
3f8590fe44
when found in CrawlURI, add WARC-Simple-Form-Province-Status header
2013-03-11 20:57:41 -07:00
gojomo
bf9cab8eb2
when POST/submit-data requested, alter request method accordingly
2013-03-11 20:57:09 -07:00
gojomo
bbac7d7320
per (overlay) settings, spawn high-priority POST when appropriate; annotate/add-headers to describe
2013-03-11 20:56:31 -07:00
gojomo
ed639c8342
extract HTMLForm INPUT details
2013-03-11 20:55:20 -07:00
gojomo
94a61b4fb7
retain offsets of discovered FORMs to aid later processing
2013-03-11 20:53:41 -07:00
gojomo
6b0abae8ee
new constants, hop-type; refactoring to help support form-submissions
2013-03-11 20:52:59 -07:00
Noah Levitt
8cbb84b8df
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-02-22 15:43:30 -08:00
Noah Levitt
fb758e53e3
Avoid pulling in hadoop-core.jar and its dependencies; add plugin versions to get rid of warnings
2013-02-22 15:43:15 -08:00
Noah Levitt
878e3b1afe
Remove spurious "throws" declaration
2013-02-05 18:37:22 -08:00
Noah Levitt
543e8820ef
Merge branch 'master' into archive-commons-refactor
2012-12-20 02:44:57 -08:00
Noah Levitt
3a65fa25d7
Working on refactoring some stuff into archive-commons
2012-12-18 14:46:52 -08:00
Noah Levitt
0eac126f7d
Special WARC-Profile for uri-agnostic revisit records. See HER-2022, http://tech.groups.yahoo.com/group/archive-crawler/message/7806
2012-12-13 14:11:14 -08:00
Noah Levitt
5586d5452d
* JerichoExtractorHTMLTest.java
...
avoid redundancy by using extractor built in setUp()
makeExtractor() - call extractor.setExtractorJS(new ExtractorJS()) since we got rid of the static ExtractorJS stuff
testConditionalComment1() - override to skip the test since it fails with JerichoExtractorHTML
2012-12-11 17:00:23 -08:00
Noah Levitt
85118c59c4
Tighten up javascript extraction based on real-world analysis (addresses HER-1523)
...
* UriUtils.java
new method isVeryLikelyUri() with tighter heuristic than isLikelyUri()
* ExtractorJS.java
use UriUtils.isVeryLikelyUri(), and change order of operations to do fixup before call to isVeryLikelyUri(), since it doesn't expect strings with javascript escaping and stuff
* StringExtractorTestBase.java
handle test data with expected value null, meaning no outlinks expected
* ExtractorHTMLTest.java
avoid redundancy by using the extractor created in ContentExtractorTestBase.setUp()
* ExtractorJSTest.java
some new tests
2012-12-11 16:27:26 -08:00
Noah Levitt
b1cd3d30a0
* ExtractorHTMLTest.java
...
makeExtractor() - call setExtractorJS(new ExtractorJS()) so testSpeculativeLinkExtraction() passes
2012-12-10 17:31:55 -08:00
Noah Levitt
256b16249a
* ExtractorJS.java
...
refactor so considerStrings() is not static, allowing it to be overridden in subclasses
* ExtractorHTML.java, ExtractorJS.java
add @Autowired parameter extractorJS used to process inline javascript, instead of call to static ExtractorJS.considerStrings()
2012-12-10 16:51:42 -08:00
Noah Levitt
e44cf9ca9b
A little cleanup of FetchHTTPTest refactoring
2012-12-04 18:43:15 -08:00
Noah Levitt
5354960f00
Refactor FetchHTTPTest to use a TestSuite so that starting and shutting down the test http servers happen only once at start and finish respectively
2012-12-04 17:25:09 -08:00
Noah Levitt
e30cb83540
* FetchHTTPTest.java
...
testLaxUrlEncoding() - Tests a URL not correctly url-encoded, but that heritrix lets pass through to mimic browser behavior.
2012-12-04 16:25:09 -08:00
Noah Levitt
92b302f164
* ExtractorMultipleRegexTest.java
...
avoid "constant string too long" compile error
2012-11-16 19:08:18 -08:00
Noah Levitt
94df568b4f
* ExtractorMultipleRegex.java
...
remove temporary performance testing code
2012-11-16 18:34:28 -08:00
Noah Levitt
aa98ec3327
* ExtractorMultipleRegex.java
...
keep cache of groovy Template objects, since they are expensive to create
2012-11-16 18:30:41 -08:00
Noah Levitt
eeee01a52d
* ExtractorMultipleRegex.java
...
javadocs
2012-11-16 16:42:43 -08:00
Noah Levitt
89dfef468c
* ExtractorMultipleRegexTest.java
...
test passes now; extractor gets what it can, which is most of the scroll down urls
2012-11-15 19:18:59 -08:00
Noah Levitt
dfd7a8a7d3
* ExtractorMultipleRegexTest.java
...
turns out that __adt parameter can be found near the json blob - most, but not all, expected links are now found
2012-11-15 18:58:11 -08:00
Noah Levitt
22ea489c57
* ExtractorMultipleRegex.java
...
remove the "fooIndex" thing from available bindings, since it's kinda hacky and turned out not to be needed for our use case
2012-11-15 18:56:54 -08:00
Noah Levitt
4af58f38e5
* CrawlURI.java
...
outLinks, outCandidates - use LinkedHashSet to ensure predictable order (any reason not to do that?)
2012-11-15 18:55:37 -08:00
Noah Levitt
75280a13a8
* ExtractorMultipleRegexTest.java
...
working on making test work - first two scroll-down urls are extracted successfully, others fail
2012-11-15 17:27:41 -08:00
Noah Levitt
13e1f9d64d
* ExtractorMultipleRegex.java
...
fix omission from last commit, uriRegex
2012-11-15 14:01:01 -08:00
Noah Levitt
6ec4b5a973
* ExtractorMultipleRegex.java
...
more refactoring to avoid redundant operations
2012-11-15 13:55:11 -08:00
Noah Levitt
43dbcbfe81
* ExtractorMultipleRegex.java
...
refactor for readability
2012-11-15 13:35:16 -08:00
Noah Levitt
3f44a4bc66
Merge branch 'master' into MoreImpliedExtractor
2012-11-15 12:35:17 -08:00
Noah Levitt
8395e372e4
* ScriptedProcessor.java
...
getEngine() - fix logic to respect isolateThreads setting
2012-11-15 12:33:48 -08:00
Noah Levitt
ba5bca807e
* ExtractorMultipleRegex.java
...
use groovy templating facility
2012-11-15 12:32:42 -08:00
Noah Levitt
6e2d2129fe
* ExtractorMultipleRegexTest.java
...
use real use case actual content, actual urls from facebook for testing
* ExtractorMultipleRegex.java
some comments, little twiddles
2012-11-15 11:16:57 -08:00
Noah Levitt
1e2d86dc63
* ContentExtractorTestBase.java
...
createRecorder(String,String) - new method for creating test Recorder and specifying charset
createRecorder(String) - deprecate (use default charset as before)
2012-11-15 11:14:25 -08:00
Noah Levitt
1d994493e2
more progress on ExtractorMultipleRegex
2012-11-14 19:58:43 -08:00