Adam Miller
646133ce4e
pom.xml
...
Adding heritrix-engine dependency
XmlCrawlSummaryReport.java
XML version of the CrawlSummaryReport
ExtractorYoutubeFormatStream.java
Adding sample bean to docs
2013-06-12 16:43:28 -07:00
Noah Levitt
1743817e1e
make code formatting consistent with main heritrix proj by changing spaces to tabs; also remove unused imports from HBase.java, no other substantive changes
2013-06-11 15:55:35 -07:00
Adam Miller
b7720ed254
Merge branch 'master' of https://github.com/internetarchive/heritrix3
2013-06-11 15:42:22 -07:00
Adam Miller
ed6cc473f8
ExtractorYoutubeFormatStream.java
...
Adjusting json extraction regex to account for semi-colons within the content
ExtractorYoutubeFormatStreamTest.java
New test case for semi-colon in json content
2013-06-11 15:40:07 -07:00
Kenji Nagahashi
9ef2fb0f7f
Merge branch 'topic/scripting-xml-fix'
2013-06-10 15:36:13 -07:00
Kenji Nagahashi
358afd4486
revised scripting console xml response fix based on comments.
...
removed scripting console actions from ScriptModel into ScriptingConsole
class (resurrection of ScriptExec).
added basic ScriptingConsoleTest.
2013-06-10 14:17:27 -07:00
Noah Levitt
f2984881f6
address HER-2042 - new annotation duplicate:uriAgnosticDigest for all url agnostic duplicates, whether written to warcs or not; stats consult this annotation so that duplicate/novel numbers are correct
2013-06-07 17:17:24 -07:00
Noah Levitt
f478798935
implement itagPriority and extractLimit parameters
2013-06-06 16:15:13 -07:00
Noah Levitt
593cae448e
test crawl reveals generate_204 urls apparently not needed
2013-06-06 14:41:26 -07:00
Noah Levitt
9a9e49d92d
only extract urls that look like youtube watch pages
2013-06-06 14:28:53 -07:00
Noah Levitt
9828310f60
make ExtractorYoutubeFormatStream extend Extractor instead of ContentExtractor because we don't really want the extra stuff that ContentExtractor does
2013-06-06 14:13:26 -07:00
Adam Miller
a7dda8b785
ExtractorYoutubeFormatStream.java
...
Parse parameters instead of attempting to regex desired values
Add priority itag list and extract limit
ExtractorYoutubeFormatStreamTest.java
Test cases for extracting all videos and extracting single video
2013-06-05 16:39:10 -07:00
Noah Levitt
11f8fd51c6
extractIndependently=true so that earlier ExtractorHTML extraction doesn't stop us from extracting here; also use java.net.URLDecoder
2013-06-05 13:25:25 -07:00
Noah Levitt
feec423b52
import Adam's ExtractorYoutubeFormatStream from its former home https://github.com/adam-miller/ExtractorYoutubeFormatStream
2013-06-05 13:16:55 -07:00
Noah Levitt
0508abb9ad
new contrib subproject for extra generally useful modules not part of the main distribution
2013-06-05 13:12:36 -07:00
Noah Levitt
ffd248f780
HER-2040 ExtractorJS shouldExtract when content-type is application/json
2013-06-03 11:23:16 -07:00
Kenji Nagahashi
3a54600a79
ADD unit test for XmlMarshaller
2013-05-29 12:03:49 -07:00
Kenji Nagahashi
c53b821c85
FIX: 73f58dc56a broke XML variant of scripting console response, and will
...
produce incompatible structure (extra nesting with scriptExec element).
moved ScriptExecution code to ScriptModel, which is now a bean, not Map.
XmlMarshaller is extended to support simple JavaBeans to XML marshalling,
borrowing annotation classes from javax.xml.
2013-05-28 13:37:04 -07:00
Kenji Nagahashi
7f2aacba2c
BeanLookupBindings: check appCtx != null explicitly, rather than using
...
assert.
2013-05-28 13:20:23 -07:00
Noah Levitt
73888819dc
fix bug where scripting console no longer retains your script when you hit execute
2013-05-15 14:23:53 -07:00
Noah Levitt
d4c5bd6e98
avoid npe by setting default <input> type "text" when unspecified
2013-05-14 18:17:09 -07:00
Noah Levitt
d972731958
ExtractorHTMLForms.java - handle a corner case extracting attribute values
2013-05-14 16:23:41 -07:00
Noah Levitt
7bd010e3c7
HTMLForm.java - default html <input> type is text, so consider <input> with no type as candidateUsernameInput
2013-05-13 22:53:37 -07:00
Noah Levitt
bea5c6f267
followup on HER-2031 ExtractorHTMLForms.java - improve regexes looking for form and input attributes, mainly to avoid maxing cpu for long periods of time for certain input
2013-05-13 22:37:02 -07:00
Noah Levitt
5b4d8c7f0d
use intended File.separator instead of File.pathSeparator
2013-05-13 22:06:09 -07:00
Noah Levitt
ec824e9724
avoid need for getRawData() method by introducing Link.hasData()
2013-05-13 17:43:45 -07:00
Kristinn Sigurðsson
7c5b2ceb00
Added a data map to Link. Becomes the CrawlURI's data map when Link is
...
promoted to CrawlURI
2013-05-13 15:18:02 +00:00
Kenji Nagahashi
3841d096d5
removed commented-out old code.
2013-05-11 15:44:25 -07:00
Kenji Nagahashi
2825cb736a
script conole: say "line" when linesExecuted==1.
...
generate global variables guide from model.availableGlobalVariables.
2013-05-11 15:44:25 -07:00
Kenji Nagahashi
73f58dc56a
scripting console:
...
fix bug: exception from script execution is not displayed.
remove extra space around exception and rawOutput.
involves a bit more than bug fix. introduced ScriptExecution class for
encapsulating script, execution, and result.
2013-05-11 15:44:25 -07:00
Noah Levitt
e531de0ed3
avoid polluting logs with long stacktraces when a request comes in for a report from an unbuilt job (a fairly common case in archive-it)
2013-04-29 19:02:22 -07:00
Noah Levitt
d5693563f8
fix bug pointed out in https://webarchive.jira.com/browse/ARI-3176?focusedCommentId=35314 - url-agnostic dedupe not reported in host report (and other reports)
2013-04-29 18:11:24 -07:00
Noah Levitt
dafd95eef7
fix for HER-2037 final crawl reports written after crawl is "FINISHED" - call dumpReports() in StatisticsTracker.stop() which (since the HER-1954 changes) runs at the right time, after all url processing has finished but before the FINISHED state is officially reached
2013-04-22 21:23:43 -07:00
Noah Levitt
1713093699
remove stray space
2013-04-22 21:18:35 -07:00
Noah Levitt
81d0a23e50
restore build timestamp
2013-04-09 10:42:23 -07:00
Noah Levitt
3d90506b99
use LinkedHashMap to ensure consistent ordering of xml api
2013-04-09 10:34:02 -07:00
Noah Levitt
08e18fd692
avoid creating unnecessary AlerthThreadGroups
2013-04-05 09:35:23 -07:00
Noah Levitt
475d07e8ca
attempt to address HER-2036 occasional strange memory leak after crawl finishes - have ToePool discard its reference to CrawlController at crawl finish
2013-04-05 09:34:58 -07:00
Noah Levitt
d07bf80e02
add link to engine page on job page
2013-04-03 23:50:23 -07:00
Noah Levitt
5a7cd8a6a8
disable codemirror search which add-on which seems to fight with browser search all the time
2013-04-03 17:16:04 -07:00
Noah Levitt
fd668e82ba
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-04-03 16:39:45 -07:00
Noah Levitt
9f33b3b667
button on engine page to ask java to run garbage collector
2013-04-03 16:39:32 -07:00
Kenji Nagahashi
c98c1d0e29
changed CrawlerLoggerModule.setupLogs() from private to protected so that sub-class can modify logging configuration.
2013-04-02 17:55:03 -07:00
Kenji Nagahashi
4dd46e6af8
quick fix to ClassCastException in ToePool.shortReportLineTo().
...
fixed a typo in Job.ftl.
2013-04-01 23:45:22 -07:00
Noah Levitt
8a40f9ff13
fix misleading logging
2013-04-01 13:04:33 -07:00
Ilya Kreymer
45b205061f
Fix for testCompressedARCStream - if arc has gzip header but doesn't have IA flags, its still compressed
2013-03-31 18:54:12 -07:00
Noah Levitt
6a0ed067c0
consolidate disparate versions of WARCConstants into one version that lives in ia-web-commons
2013-03-29 17:12:32 -07:00
Noah Levitt
1ac1dd9bab
fix typo which removed call to candidate.setFullVia(), should fix failing Precedence3SelfTest
2013-03-29 13:17:58 -07:00
Noah Levitt
dcbbec4bbb
fix for HER-2035 xml representation of job url can be invalid
2013-03-29 12:16:07 -07:00
Noah Levitt
5fdf5713a2
Merge branch 'master' of github.com:internetarchive/heritrix3
2013-03-28 19:10:19 -07:00