Commit Graph
1109 Commits
Author SHA1 Message Date
Noah Levitt 52d72e661d part of HER-2048 - WARCWriterProcessor stats resume from checkpoint 2013-09-07 11:35:03 -07:00
Noah Levitt e0f78976d7 part of HER-2048 - restore BdbServerCache from checkpoint 2013-09-07 10:53:22 -07:00
Noah Levitt fd71833050 avoid NPE when attempting to terminate already-finished job 2013-09-07 09:57:33 -07:00
Noah Levitt d01bc4ed55 when resuming from checkpoint, make next checkpoint serial number follow serial number of resumed checkpoint 2013-09-07 09:57:07 -07:00
Noah Levitt d61171f076 CrawlController.reserveMemory - use byte[] instead of char[] for clarity of size 2013-09-07 09:55:58 -07:00
Noah Levitt b670a3c2a7 Merge branch 'master' into always-resumable 2013-09-06 23:32:32 -07:00
Noah Levitt 25d4de186f don't log warning in normal case 2013-09-06 22:04:56 -07:00
Noah Levitt cc0b63066b Merge branch 'master' into always-resumable 2013-09-03 18:38:55 -07:00
Noah Levitt 34cb015ecf adjust PublicSuffixesTest for updated public suffixes list 2013-09-03 18:38:44 -07:00
Noah Levitt 59596c7043 prevent redundant checkpoints when crawl is stopping 2013-09-03 14:24:41 -07:00
Noah Levitt 0ab8eee24e forgetting old checkpoints - roll up previous checkpointed CrawlerJournals (e.g. frontier.recover.gz) into new checkpoint 2013-08-29 10:56:12 -07:00
Noah Levitt 73f6b20e48 Merge branch 'master' into always-resumable 2013-08-28 18:31:42 -07:00
Noah Levitt 981cf5bc9d update public suffixes list 2013-08-28 18:31:28 -07:00
Noah Levitt 1315717002 Merge branch 'master' into always-resumable 2013-08-28 18:14:03 -07:00
Noah Levitt 418c7138a9 handle case where there is no whois server for a domain 2013-08-28 18:13:53 -07:00
Noah Levitt e8c5083d1a handle case where there is no whois server for a domain 2013-08-28 18:13:33 -07:00
Noah Levitt 43bb42d346 BdbModule#doCheckpoint - support forgetting all but latest checkpoint 2013-08-28 11:22:37 -07:00
Noah Levitt 8e7c26bc8f fix setForgetAllButLatest 2013-08-28 10:58:32 -07:00
Noah Levitt 2fe7d8cd45 some progress on forgetting old checkpoints when checkpointing anew 2013-08-27 18:39:25 -07:00
Noah Levitt 77a12e63b8 make special "host" "whois:" appear correctly in hosts report 2013-08-27 17:39:07 -07:00
Noah Levitt 4abfcb01b0 HER-1895 support HTTP Refresh header 2013-08-26 18:33:02 -07:00
Noah Levitt d1bb1a4637 further improve performance of BdbUriUniqFilter.forgetAllSchemeAuthorityMatching() by skipping conversion of key data to long and comparing bytes directly 2013-08-21 11:44:46 -07:00
Noah Levitt a06c141740 rename BdbUriUniqFilter.forgetSchemeHost() to forgetAllSchemeAuthorityMatching(), improve performance 2013-08-21 11:31:05 -07:00
Noah Levitt a6958951df Reset Recorder state when uri processing is finished (had been seeing cases where a url's entry in crawl.log would sometimes have nonzero size value if the url was erroring out and RecordingOutputStream.open() was never called) 2013-08-12 19:28:28 -07:00
Noah Levitt 68adaa70fd need to close the recorder here so it can calculate the content length properly 2013-08-09 19:14:58 -07:00
Noah Levitt 705a375daf uses of UriUtils.isLikelyUri() in Extractor{HTML,SWF,XML} with UriUtils.isVeryLikelyUri() to reap the benefits of HER-1523 improvements (should address archive-it issue ARI-3492) 2013-08-09 18:17:11 -07:00
Noah Levitt 8c64cba9ae TransclusionDecideRule.java - consider form submission "S" hop as navlink like "L", not transcluded 2013-08-05 16:05:02 -07:00
Noah Levitt f4f3b9be38 add logging to BdbUriUniqFilter.forgetSchemeHost(), clean up some javadocs 2013-07-26 17:50:36 -07:00
Noah Levitt b9726c8d4f new method BdbUriUniqFilter.forgetSchemeHost(String schemeHost), and unit test 2013-07-25 18:51:26 -07:00
Noah Levitt b6b3ae6970 make log level check for debug/performance logging consistent to avoid unnecessary operations 2013-07-25 17:28:16 -07:00
Noah Levitt 72094213a4 MirrorWriterProcessor.java - change "path" setting to a ConfigPath and set default value "${launchId}/mirror" 2013-07-22 17:54:06 -07:00
Ilya Kreymer d37b43e4f3 Merge branch 'master' of github.com:internetarchive/heritrix3 2013-07-18 21:01:10 -07:00
Ilya Kreymer b60c43fb81 ARCRecord: Add error offset to position when adjusting for bad header lines 2013-07-18 21:00:01 -07:00
Noah Levitt 452304ef12 profile-crawler-beans.cxml - include commented out default value for WARCWriterProcessor "template" setting - see http://tech.groups.yahoo.com/group/archive-crawler/message/8182 2013-07-18 18:38:33 -07:00
Ilya Kreymer 0c555b4b62 Merge branch 'master' of github.com:internetarchive/heritrix3 2013-07-18 17:11:08 -07:00
Ilya Kreymer 1b67f5ce6b ArcRecord: Silently skip any header fields before a status line in arc record. This allows the arc reader to be fault-tolerant of 'bad' arcs that have extra headers
inserted before the status line (but are otherwise valid)
2013-07-18 17:08:38 -07:00
Adam Miller dd135c9383 Making youtube extractor log messages less noisy. 2013-06-25 11:33:24 -07:00
Ilya Kreymer 3916175e4d FIX: Add ArchiveRecordHeader.getContentLength() to specifically return the warc/arc record content length,
not including headers
2013-06-21 14:56:31 -07:00
Adam Miller b85828135b Convert MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule to use CrawlURI.getFetchStatus() to avoid null status codes and broaden use.
Added test cases for MatchesStatusCodeDecideRule and NotMatchesStatusCodeDecideRule
Fix comments for youtube extractor.
2013-06-21 12:03:41 -07:00
Kenji Nagahashi 1e3e66ac4c SingleHBaseTable: FIX error message typo, print friendlier message for
TableNotFoundException.
2013-06-16 08:23:12 -07:00
Kenji Nagahashi b88242ee16 bean browser: handle array of non-Object element type properly.
(just avoiding throwing exception - further change would be necessary
for better rendering.)
2013-06-16 08:21:53 -07:00
Kenji Nagahashi 030006ddd5 drop debug="true" from log4j.xml in contrib 2013-06-13 18:45:55 -07:00
Kenji Nagahashi 9bf098fd44 add contrib/target to .gitignore 2013-06-13 18:43:35 -07:00
Kenji Nagahashi b2a0495a08 merged Kenji's change to hbase de-duplication module:
- SingleHBaseTable using single HTable instances (no HTablePool) with
backing-off upon communication errors.
- FetchHistoryHelper for allowing multiple crawl history sources.
- row key backward compatibility mode.
2013-06-13 17:21:28 -07:00
Adam Miller 646133ce4e pom.xml
Adding heritrix-engine dependency

XmlCrawlSummaryReport.java
	XML version of the CrawlSummaryReport

ExtractorYoutubeFormatStream.java
	Adding sample bean to docs
2013-06-12 16:43:28 -07:00
Noah Levitt 1743817e1e make code formatting consistent with main heritrix proj by changing spaces to tabs; also remove unused imports from HBase.java, no other substantive changes 2013-06-11 15:55:35 -07:00
Adam Miller b7720ed254 Merge branch 'master' of https://github.com/internetarchive/heritrix3 2013-06-11 15:42:22 -07:00
Adam Miller ed6cc473f8 ExtractorYoutubeFormatStream.java
Adjusting json extraction regex to account for semi-colons within the content

ExtractorYoutubeFormatStreamTest.java
	New test case for semi-colon in json content
2013-06-11 15:40:07 -07:00
Kenji Nagahashi 9ef2fb0f7f Merge branch 'topic/scripting-xml-fix' 2013-06-10 15:36:13 -07:00
Kenji Nagahashi 358afd4486 revised scripting console xml response fix based on comments.
removed scripting console actions from ScriptModel into ScriptingConsole
class (resurrection of ScriptExec).
added basic ScriptingConsoleTest.
2013-06-10 14:17:27 -07:00