John MacFarlaneandClaude 6d7e2f6544 Text.Pandoc.XML.Light: parse XML fragments from the event stream.
parseXMLContentsWithEntities previously used xml-conduit's document
parser, which requires a single root element, and worked around this
by re-parsing fragments wrapped in "<wrapper>...</wrapper>" when the
first parse failed with ContentAfterRoot.  This was fragile: it broke
on fragments with an XML declaration or DOCTYPE before multiple root
elements, failed on some text-only input, parsed fragments twice, and
gave different results depending on which path was taken.

We now drive xml-conduit's streaming parser directly, folding the
event stream into a Content forest in a single pass, with start/end
tag balance checked during the fold (the stream parser itself does
not validate this).  Behavior changes:

- Content fragments with an XML declaration or DOCTYPE followed by
  multiple root elements, and text-only or empty input, now parse
  instead of erroring.
- Attributes now preserve document order instead of being sorted
  alphabetically (the document parser stored them in a Map).
- Errors for unresolved entities and mismatched tags now report
  source positions.

This is also slightly faster than the old document-parser path.
parseXMLElementWithEntities is unchanged, still requiring a single
root element.  The xml-light library component gains a direct
dependency on conduit (already a transitive dependency via
xml-conduit).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 22:49:09 +00:00
2026-09-06 10:39:38 -07:00
…
2026-07-25 13:07:30 +02:00
…
2026-08-28 10:11:37 -07:00
…
2026-08-28 11:04:16 -07:00
2026-07-12 12:09:45 +02:00
…
2026-07-24 00:24:01 +02:00
…

Pandoc

github
release hackage
release homebrew stackage LTS
package CI
tests license pandoc-discuss on google
groups

The universal markup converter

Pandoc is a Haskell library for converting from one markup format to another, and a command-line tool that uses this library.

It can convert from

It can convert to

Pandoc can also produce PDF output via LaTeX, Groff ms, or HTML.

Pandoc’s enhanced version of Markdown includes syntax for tables, definition lists, metadata blocks, footnotes, citations, math, and much more. See the User’s Manual below under Pandoc’s Markdown.

Pandoc has a modular design: it consists of a set of readers, which parse text in a given format and produce a native representation of the document (an abstract syntax tree or AST), and a set of writers, which convert this native representation into a target format. Thus, adding an input or output format requires only adding a reader or writer. Users can also run custom pandoc filters to modify the intermediate AST (see the documentation for filters and Lua filters).

Because pandoc’s intermediate representation of a document is less expressive than many of the formats it converts between, one should not expect perfect conversions between every format and every other. Pandoc attempts to preserve the structural elements of a document, but not formatting details such as margin size. And some document elements, such as complex tables, may not fit into pandoc’s simple document model. While conversions from pandoc’s Markdown to all formats aspire to be perfect, conversions from formats more expressive than pandoc’s Markdown can be expected to be lossy.

Installing

Here’s how to install pandoc.

Documentation

Pandoc’s website contains a full User’s Guide. It is also available here as pandoc-flavored Markdown. The website also contains some examples of the use of pandoc, a limited online demo, and a WebAssembly-based online demo.

Contributing

Pull requests, bug reports, and feature requests are welcome. Please make sure to read the contributor guidelines before opening a new issue.

License

© 2006-2024 John MacFarlane (jgm@berkeley.edu). Released under the GPL, version 2 or greater. This software carries no warranty of any kind. (See COPYRIGHT for full copyright and warranty notices.)

Languages
Haskell 82.1%
Roff 6%
Rich Text Format 4.7%
HTML 2.1%
Lua 1.9%
Other 2.9%