diff --git a/dist/README.txt b/dist/README.txt index 8a88c23d..c96bf635 100644 --- a/dist/README.txt +++ b/dist/README.txt @@ -9,8 +9,9 @@ Readme for Heritrix 6. License -1. Introduction +1) Introduction ---------------- + Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project. Heritrix (sometimes spelled heretrix, or misspelled or missaid as heratrix/heritix/heretix/heratix) is an archaic word @@ -19,8 +20,9 @@ preserve the digital artifacts of our culture for the benefit of future researchers and generations, this name seemed apt. -2. Crawl Operators! +2) Crawl Operators! -------------------- + Heritrix is designed to respect the robots.txt exclusion directives and META robots tags . Please consider the @@ -30,25 +32,29 @@ User-Agent so sites that may be adversely affected by your crawl can contact you or adapt their server behavior accordingly. -3. Getting Started +3) Getting Started ------------------- -See the User Manual at + +See the User Manual, available from . -For API documentation, see +For API documentation, see and -5. Latest Releases +5) Latest Releases ------------------- + Information about releases can be found at -6. License +6) License ----------- + Heritrix is free software; you can redistribute it and/or modify it under the terms of the Apache License, Version 2.0: