Noah Levitt
0ab8eee24e
forgetting old checkpoints - roll up previous checkpointed CrawlerJournals (e.g. frontier.recover.gz) into new checkpoint
2013-08-29 10:56:12 -07:00
Noah Levitt
73f6b20e48
Merge branch 'master' into always-resumable
2013-08-28 18:31:42 -07:00
Noah Levitt
981cf5bc9d
update public suffixes list
2013-08-28 18:31:28 -07:00
Noah Levitt
1315717002
Merge branch 'master' into always-resumable
2013-08-28 18:14:03 -07:00
Noah Levitt
418c7138a9
handle case where there is no whois server for a domain
2013-08-28 18:13:53 -07:00
Noah Levitt
e8c5083d1a
handle case where there is no whois server for a domain
2013-08-28 18:13:33 -07:00
Noah Levitt
43bb42d346
BdbModule#doCheckpoint - support forgetting all but latest checkpoint
2013-08-28 11:22:37 -07:00
Noah Levitt
8e7c26bc8f
fix setForgetAllButLatest
2013-08-28 10:58:32 -07:00
Noah Levitt
2fe7d8cd45
some progress on forgetting old checkpoints when checkpointing anew
2013-08-27 18:39:25 -07:00
Noah Levitt
77a12e63b8
make special "host" "whois:" appear correctly in hosts report
2013-08-27 17:39:07 -07:00
Noah Levitt
4abfcb01b0
HER-1895 support HTTP Refresh header
2013-08-26 18:33:02 -07:00
Noah Levitt
d1bb1a4637
further improve performance of BdbUriUniqFilter.forgetAllSchemeAuthorityMatching() by skipping conversion of key data to long and comparing bytes directly
2013-08-21 11:44:46 -07:00
Noah Levitt
a06c141740
rename BdbUriUniqFilter.forgetSchemeHost() to forgetAllSchemeAuthorityMatching(), improve performance
2013-08-21 11:31:05 -07:00
Noah Levitt
a6958951df
Reset Recorder state when uri processing is finished (had been seeing cases where a url's entry in crawl.log would sometimes have nonzero size value if the url was erroring out and RecordingOutputStream.open() was never called)
2013-08-12 19:28:28 -07:00
Noah Levitt
68adaa70fd
need to close the recorder here so it can calculate the content length properly
2013-08-09 19:14:58 -07:00
Noah Levitt
705a375daf
uses of UriUtils.isLikelyUri() in Extractor{HTML,SWF,XML} with UriUtils.isVeryLikelyUri() to reap the benefits of HER-1523 improvements (should address archive-it issue ARI-3492)
2013-08-09 18:17:11 -07:00
Noah Levitt
8c64cba9ae
TransclusionDecideRule.java - consider form submission "S" hop as navlink like "L", not transcluded
2013-08-05 16:05:02 -07:00
Noah Levitt
f4f3b9be38
add logging to BdbUriUniqFilter.forgetSchemeHost(), clean up some javadocs
2013-07-26 17:50:36 -07:00
Noah Levitt
b9726c8d4f
new method BdbUriUniqFilter.forgetSchemeHost(String schemeHost), and unit test
2013-07-25 18:51:26 -07:00
Noah Levitt
b6b3ae6970
make log level check for debug/performance logging consistent to avoid unnecessary operations
2013-07-25 17:28:16 -07:00
Noah Levitt
72094213a4
MirrorWriterProcessor.java - change "path" setting to a ConfigPath and set default value "${launchId}/mirror"
2013-07-22 17:54:06 -07:00
Ilya Kreymer
d37b43e4f3
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-07-18 21:01:10 -07:00
Ilya Kreymer
b60c43fb81
ARCRecord: Add error offset to position when adjusting for bad header lines
2013-07-18 21:00:01 -07:00
Noah Levitt
452304ef12
profile-crawler-beans.cxml - include commented out default value for WARCWriterProcessor "template" setting - see http://tech.groups.yahoo.com/group/archive-crawler/message/8182
2013-07-18 18:38:33 -07:00
Ilya Kreymer
0c555b4b62
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-07-18 17:11:08 -07:00
Ilya Kreymer
1b67f5ce6b
ArcRecord: Silently skip any header fields before a status line in arc record. This allows the arc reader to be fault-tolerant of 'bad' arcs that have extra headers
...
inserted before the status line (but are otherwise valid)
2013-07-18 17:08:38 -07:00
Adam Miller
dd135c9383
Making youtube extractor log messages less noisy.
2013-06-25 11:33:24 -07:00
Ilya Kreymer
3916175e4d
FIX: Add ArchiveRecordHeader.getContentLength() to specifically return the warc/arc record content length,
...
not including headers
2013-06-21 14:56:31 -07:00
Adam Miller
b85828135b
Convert MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule to use CrawlURI.getFetchStatus() to avoid null status codes and broaden use.
...
Added test cases for MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule
Fix comments for youtube extractor.
2013-06-21 12:03:41 -07:00
Kenji Nagahashi
1e3e66ac4c
SingleHBaseTable: FIX error message typo, print friendlier message for
...
TableNotFoundException.
2013-06-16 08:23:12 -07:00
Kenji Nagahashi
b88242ee16
bean browser: handle array of non-Object element type properly.
...
(just avoiding throwing exception - further change would be necessary
for better rendering.)
2013-06-16 08:21:53 -07:00
Kenji Nagahashi
030006ddd5
drop debug="true" from log4j.xml in contrib
2013-06-13 18:45:55 -07:00
Kenji Nagahashi
9bf098fd44
add contrib/target to .gitignore
2013-06-13 18:43:35 -07:00
Kenji Nagahashi
b2a0495a08
merged Kenji's change to hbase de-duplication module:
...
- SingleHBaseTable using single HTable instances (no HTablePool) with
backing-off upon communication errors.
- FetchHistoryHelper for allowing multiple crawl history sources.
- row key backward compatibility mode.
2013-06-13 17:21:28 -07:00
Adam Miller
646133ce4e
pom.xml
...
Adding heritrix-engine dependency
XmlCrawlSummaryReport.java
XML version of the CrawlSummaryReport
ExtractorYoutubeFormatStream.java
Adding sample bean to docs
2013-06-12 16:43:28 -07:00
Noah Levitt
1743817e1e
make code formatting consistent with main heritrix proj by changing spaces to tabs; also remove unused imports from HBase.java, no other substantive changes
2013-06-11 15:55:35 -07:00
Adam Miller
b7720ed254
Merge branch 'master' of https://github.com/internetarchive/heritrix3
2013-06-11 15:42:22 -07:00
Adam Miller
ed6cc473f8
ExtractorYoutubeFormatStream.java
...
Adjusting json extraction regex to account for semi-colons within the content
ExtractorYoutubeFormatStreamTest.java
New test case for semi-colon in json content
2013-06-11 15:40:07 -07:00
Kenji Nagahashi
9ef2fb0f7f
Merge branch 'topic/scripting-xml-fix'
2013-06-10 15:36:13 -07:00
Kenji Nagahashi
358afd4486
revised scripting console xml response fix based on comments.
...
removed scripting console actions from ScriptModel into ScriptingConsole
class (resurrection of ScriptExec).
added basic ScriptingConsoleTest.
2013-06-10 14:17:27 -07:00
Noah Levitt
f2984881f6
address HER-2042 - new annotation duplicate:uriAgnosticDigest for all url agnostic duplicates, whether written to warcs or not; stats consult this annotation so that duplicate/novel numbers are correct
2013-06-07 17:17:24 -07:00
Noah Levitt
f478798935
implement itagPriority and extractLimit parameters
2013-06-06 16:15:13 -07:00
Noah Levitt
593cae448e
test crawl reveals generate_204 urls apparently not needed
2013-06-06 14:41:26 -07:00
Noah Levitt
9a9e49d92d
only extract urls that look like youtube watch pages
2013-06-06 14:28:53 -07:00
Noah Levitt
9828310f60
make ExtractorYoutubeFormatStream extend Extractor instead of ContentExtractor because we don't really want the extra stuff that ContentExtractor does
2013-06-06 14:13:26 -07:00
Adam Miller
a7dda8b785
ExtractorYoutubeFormatStream.java
...
Parse parameters instead of attempting to regex desired values
Add priority itag list and extract limit
ExtractorYoutubeFormatStreamTest.java
Test cases for extracting all videos and extracting single video
2013-06-05 16:39:10 -07:00
Noah Levitt
11f8fd51c6
extractIndependently=true so that earlier ExtractorHTML extraction doesn't stop us from extracting here; also use java.net.URLDecoder
2013-06-05 13:25:25 -07:00
Noah Levitt
feec423b52
import Adam's ExtractorYoutubeFormatStream from its former home https://github.com/adam-miller/ExtractorYoutubeFormatStream
2013-06-05 13:16:55 -07:00
Noah Levitt
0508abb9ad
new contrib subproject for extra generally useful modules not part of the main distribution
2013-06-05 13:12:36 -07:00
Noah Levitt
ffd248f780
HER-2040 ExtractorJS shouldExtract when content-type is application/json
2013-06-03 11:23:16 -07:00