Fix for [HER-1268] ExtractorHTML not able to extract some links (due to them violating RFC2396)

* ExtractorHTMLTest.java, UURIFactoryTest.java
    test that relative-URIs with late-position colons aren't interpreted as absolute URIs with long, illegal schemes
* UURIFactory.java
    update RFC2396REGEX to require legal scheme
* LaxURI.java
    allow illegal candidate scheme to just be interpreted as something else
This commit is contained in:
gojomo committed 2009-05-12 22:22:11 +00:00
1 parent 252cf63863
commit 2dec5dfd1a
4 files changed
+66 -8

No files matched your search

@@ -778,6 +778,25 @@ public class UURIFactoryTest extends TestCase {
"http://www.example.com/path/:foo");
}
/**
* Ensure that relative URIs with colons in late positions
* aren't mistakenly interpreted as absolute URIs with long,
* illegal schemes.
*
* @throws URIException
*/
public void testLateColon() throws URIException {
UURI base = UURIFactory.getInstance("http://www.example.com/path/page");
UURI uuri1 = UURIFactory.getInstance(base,"example.html;jsessionid=deadbeef:deadbeed?parameter=this:value");
assertEquals("derelativize lateColon",
uuri1.getURI(),
"http://www.example.com/path/example.html;jsessionid=deadbeef:deadbeed?parameter=this:value");
UURI uuri2 = UURIFactory.getInstance(base,"example.html?parameter=this:value");
assertEquals("derelativize lateColon",
uuri2.getURI(),
"http://www.example.com/path/example.html?parameter=this:value");
}
/**
* Ensure that stray trailing '%' characters do not prevent
* UURI instances from being created, and are reasonably