Commit Graph
19350 Commits
Author SHA1 Message Date
John MacFarlane 7cf4d41fe5 Use compactifyTable for tables produced by all readers.
Remove old ad hoc `paraToPlain` at the table cell level.

This should ensure that we don't get tables that mix
Plain and Para. Such tables tend to look funny when
rendered in docx and other formats.

Closes #11864.
2026-09-12 20:02:41 -07:00
John MacFarlane d4cddb5539 T.P.Shared: add new function compactifyTable.
[API change]

This converts cells that consist in a single Para block
to a Plain, provided the table contains ONLY such cells
(or empty cells).
2026-09-12 19:59:45 -07:00
John MacFarlane c9a9a5eed7 Docx writer: include default style even if paragraph has non-style props.
Previously a paragraph that was, e.g. center-aligned would be
missing its default Body Text style.

Also ensure that any block sequence (list items, table cells,
block quote, div), First Paragraph is set for the first paragraph
in the item.

Closes #11867.

further tweaks
2026-09-12 11:10:46 -07:00
John MacFarlane 6c335dc4c9 diff-golden-tests.sh - print filename 2026-09-12 11:03:43 -07:00
John MacFarlane 88d84eedfd Fix diff-zip.sh on non-Darwin. 2026-09-12 10:40:34 -07:00
John MacFarlane 59a409a5d5 Add tools/diff-golden-tests.sh. 2026-09-12 10:30:08 -07:00
John MacFarlane ae45d56461 Use latest djoths. 2026-09-11 14:04:33 -07:00
John MacFarlane 226c722359 Org reader: recognize .pdf as an image format.
This matches the behavior of Emacs, which will render
`[[file:foo.pdf]]` as an image.

See #11859.
2026-09-11 08:18:46 -07:00
John MacFarlane 0e1930aced rebase_relative_paths: recognize URLs with unknown schemes.
Closes #11858.
2026-09-10 22:14:26 -07:00
John MacFarlane 85b46569ff Use latest asciidoc-hs (more perf improvements). 2026-09-10 20:19:39 -07:00
John MacFarlane b65040f0e7 Fix a bug in jats-reader.xml (duplicate attribute) 2026-09-10 17:06:11 -07:00
John MacFarlaneandClaude 6d7e2f6544 Text.Pandoc.XML.Light: parse XML fragments from the event stream.
parseXMLContentsWithEntities previously used xml-conduit's document
parser, which requires a single root element, and worked around this
by re-parsing fragments wrapped in "<wrapper>...</wrapper>" when the
first parse failed with ContentAfterRoot.  This was fragile: it broke
on fragments with an XML declaration or DOCTYPE before multiple root
elements, failed on some text-only input, parsed fragments twice, and
gave different results depending on which path was taken.

We now drive xml-conduit's streaming parser directly, folding the
event stream into a Content forest in a single pass, with start/end
tag balance checked during the fold (the stream parser itself does
not validate this).  Behavior changes:

- Content fragments with an XML declaration or DOCTYPE followed by
  multiple root elements, and text-only or empty input, now parse
  instead of erroring.
- Attributes now preserve document order instead of being sorted
  alphabetically (the document parser stored them in a Map).
- Errors for unresolved entities and mismatched tags now report
  source positions.

This is also slightly faster than the old document-parser path.
parseXMLElementWithEntities is unchanged, still requiring a single
root element.  The xml-light library component gains a direct
dependency on conduit (already a transitive dependency via
xml-conduit).

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 22:49:09 +00:00
John MacFarlane 44b869c2fa Use latest dev asciidoc-hs (performance improvements). 2026-09-10 15:30:26 -07:00
John MacFarlane b46a3b0ef7 Use latest dev asciidoc-hs. 2026-09-10 12:26:48 -07:00
John MacFarlaneandClaude 2b68672c33 Text.Pandoc.XML.Light: avoid round-trips in ppCDataS prettify path.
Previously the prettified CData path rendered the escaped text to a
Builder, materialized it as lazy then strict Text, unpacked it to
String, and re-folded it into a Builder one singleton at a time.
Since escaping neither adds nor removes newlines, we can instead
split the unescaped text on newlines, escape each line, and
interleave the indentation, emitting whole chunks.

This makes showCData and ppcCData unused, so they are removed.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude 1926ec0e83 Text.Pandoc.XML.Light: make escStr more efficient.
Previously, if a text contained any escapable character, the
whole text was unpacked to String and rebuilt one Builder
singleton per character.  Now we copy whole chunks between
escapable characters with fromText, and avoid the extra
initial T.any pass.  About 5x faster on typical text with
occasional escapes; output is unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude f6352e497c Text.Pandoc.XML.Light: fix escaping of repeated ]]> in CDATA.
escCData stopped processing after the first "]]>" occurrence,
emitting the rest of the text verbatim; a second "]]>" would
prematurely terminate the CDATA section, producing invalid XML.
(The original xml-light code recurses after each occurrence; the
recursion was lost in the Text port.)  Rewritten with T.breakOn,
which also avoids building the output one character at a time.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude ef442564af Text.Pandoc.Readers.Muse: check raw line before parsing table rows.
Table row parsers are speculatively tried at every paragraph line
boundary via the terminator continuation.  Each of the three row
variants would fully inline-parse the line before failing for lack of
a cell separator, so every paragraph line was inline-parsed several
times.

Since cell separators (whitespace followed by pipes) cannot span
lines, scanning the raw line for one lets us reject non-table lines
without any inline parsing.  The check is conservative and never
rejects a line a row parser could accept.

Together with the previous commit, this makes the reader 4.3x faster
on a 1.6 MB test document and reduces heap allocation by 87%.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude 70910f4438 Text.Pandoc.Readers.Muse: try str earlier in inline parser.
Plain words previously fell through ~20 failing alternatives (emphasis,
tag, and link parsers) before reaching str. Since every other
alternative starts with a non-alphanumeric character and str only
matches alphanumerics, trying str right after whitespace cannot change
parse results but avoids the failed attempts for every word.

On a 1.6 MB test document this makes the reader 2.2x faster and
reduces heap allocation by 61%.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude bf52badec6 Text.Pandoc.MediaBag: prevent mediaPath collisions between keys.
The friendly mediaPath was derived by percent-unescaping the key, so
distinct keys like "a%20b.png" and "a b.png" produced the same
mediaPath ("a b.png") and silently clobbered each other on
extraction (and inside docx/epub archives). Now the original name
is only kept if the key contains no percent sign, so mediaPath
equals the key and distinct keys yield distinct paths; anything
percent-encoded gets a content-hash name. Hashed names can only
coincide for identical contents, which is harmless.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude 9a39bca15c Text.Pandoc.MediaBag: collapse . and .. components in canonicalize.
normalise does not remove redundant path components, so img/../a.png
and a.png were distinct keys: the same resource could be stored
twice, and a lookup by one spelling missed an insert by the other.
Use makeCanonical (as PandocPure's FileTree already does for its
path-indexed map), which also handles duplicate and trailing
slashes, replacing backslashes with slashes first. As a side
effect, paths with a collapsible ".." now keep a friendly name
instead of being renamed to a content hash.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude 2d9423c186 Text.Pandoc.MediaBag: only reject ".." as a path component.
The insertMedia check used isInfixOf, so a harmless name like
foo..bar.png was silently renamed to its content hash. Check for an
actual ".." path component instead, treating both / and \ as
separators since the unescaped path may contain backslashes from
percent-encoding. Added tests.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude 0675048ad9 Text.Pandoc.MediaBag: use Text.Pandoc.URI.isURI in canonicalize.
Network.URI.isURI treats Windows drive-letter paths like c:/foo.png
as URIs (the drive letter parses as a scheme), so they escaped path
normalization; and it rejects URIs containing non-ASCII characters,
so e.g. https://example.com/résumé.png fell into the file-path
branch, where normalise collapsed the double slash. Pandoc's own
isURI restricts to known schemes, escapes non-ASCII characters
before parsing, and has a fast path for base64 data URIs.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude bc47dca0ca Text.Pandoc.MediaBag: make MediaItem mime type and path strict.
Use strict fields for mediaMimeType and mediaPath, and switch to
Data.Map.Strict so that insertMedia forces them at insert time.
Previously the lazy map stored an unevaluated record construction
that retained the original path, the parseURI result, and the
mime-type fallback machinery until first lookup. mediaContents is
left lazy so contents need not be forced at insert time.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude ac4abe44e3 Treat data: and file: URI schemes case-insensitively.
URI schemes are case-insensitive (RFC 3986, section 3.1), but several
places matched only lowercase "data:" or "file:":

- Text.Pandoc.URI.pBase64DataURI (also used by isURI)
- downloadOrRead: the fast-path prefix check and the parseURI
  fallback for data: and file: URIs
- Text.Pandoc.App.Input.readSource: file:, http:, and https: URIs
  given as input sources
- Text.Pandoc.MediaBag: canonicalize's fast path and insertMedia's
  data-URI branch

Previously, e.g. DATA:image/gif;base64,... would be handed to openURL
and the image dropped, and a large uppercase data URI would also force
a full parseURI in MediaBag.canonicalize. Now such URIs are handled
the same as their lowercase equivalents. Added a test case.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlaneandClaude 23bc02c63c Text.Pandoc.MediaBag: use hashlazy to avoid copying media contents.
Hash the lazy bytestring chunk-wise instead of forcing the entire
contents into a single strict buffer with BL.toStrict. Digests are
unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 23:30:32 -07:00
John MacFarlane 693570b866 T.P.Class: logOutput: avoid multiple hPutStrLn.
They could result in a single log message's lines being
interleaved in output.
2026-09-09 13:21:22 -07:00
John MacFarlane a0595b8c2b PandocMonad: Fix typos and haddock comment positions. 2026-09-09 13:12:17 -07:00
John MacFarlaneandClaude d8ecd73767 Finish factoring openURL into Text.Pandoc.Class.IO.HTTP.
Commit 455bea907 added the new module but did not register it in
pandoc.cabal or remove the original definitions, leaving the file
uncompiled and Text.Pandoc.Class.IO with duplicate copies of
openURL and getManager.  Add the module to other-modules and make
Text.Pandoc.Class.IO re-export openURL from it, deleting the
duplicated code and now-unneeded imports.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 19:15:03 +00:00
John MacFarlaneandClaude c9d5e8ce14 Text.Pandoc.Class.PandocPure: don't follow symlink cycles in addToFileTree.
The recursive directory traversal followed symlinks without any
cycle detection, so a symlink pointing to an ancestor directory
caused non-termination.  Track the canonical paths of ancestor
directories and skip a directory that is already on the current
chain.  Symlinks to directories outside the chain are still
followed as before.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 19:11:53 +00:00
John MacFarlaneandClaude 687a437e38 Text.Pandoc.Class.PandocMonad: reset HTTP manager in setNoCheckCertificate.
The HTTP manager is created lazily with TLS settings based on
stNoCheckCertificate and then cached in CommonState, so changing the
option after the first request had no effect.  Discard the cached
manager when the option's value changes.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 19:06:20 +00:00
John MacFarlaneandClaude 1dc9cd7d22 Text.Pandoc.Class.PandocMonad: report accurate offset in toTextM errors.
The offset in PandocUTF8DecodingError was computed with B.elemIndex,
i.e. the first occurrence of the offending byte *value* anywhere in
the file.  If the same byte occurred earlier as part of a valid
multi-byte sequence, the reported position pointed at valid text.
Scan for the first invalid UTF-8 sequence instead and report its
actual position and byte.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 19:05:05 +00:00
John MacFarlaneandClaude cb749b61e6 Text.Pandoc.Class.PandocMonad: fix percent-decoding in extractURIData.
Percent-escapes in a data URI represent raw octets (RFC 2397), but
the data was decoded with unEscapeString (which UTF-8-decodes
consecutive escapes into Chars) and then re-encoded with
UTF8.fromString.  This corrupted percent-encoded binary data:
e.g. %89 became the two bytes C2 89.  Decode the escapes directly
to bytes instead.  Also split off the media type before unescaping,
so an escaped comma cannot shift the separator.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 18:55:20 +00:00
John MacFarlaneandClaude 1faed87733 Text.Pandoc.Class.PandocMonad: fix base64 detection in extractURIData.
The base64 indicator in a data URI is the final parameter of the
media type and may follow other parameters, e.g.
data:text/plain;charset=utf-8;base64,...  Previously such URIs were
not base64-decoded, because the code expected ";base64" to be the
only parameter.  The media type parameters are now also retained in
the returned MIME type, as they already are in pBase64DataURI.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 18:45:11 +00:00
John MacFarlaneandClaude 1c841b85e8 Text.Pandoc.Class.PandocMonad: fix isRelativeToParentDir.
The check used a string prefix comparison, so any path beginning
with two dots (e.g. `..foo/bar.yaml`) was treated as pointing
outside the working directory, wrongly disabling the user data
directory fallback in checkUserDataDir.  Compare the first path
component instead.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 18:36:36 +00:00
John MacFarlaneandClaude 46e7475158 Text.Pandoc.Class.PandocMonad: make runSilently error-safe.
Previously, if the action passed to runSilently threw an error that
was later caught, the verbosity remained pinned at ERROR and all
previously accumulated log messages were lost.  Now the original
log and verbosity are restored even when the action fails.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 18:34:37 +00:00
John MacFarlaneandClaude cbbbf312ea Text.Pandoc.Class.PandocMonad: avoid copying input in toTextM.
Skip the CR-filtering copy when the input contains no CRs,
as already done in Text.Pandoc.UTF8.toText.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 17:38:41 +00:00
John MacFarlaneandClaude 762d6384b2 Text.Pandoc.UTF8: make readFile exception-safe.
Use withFile instead of openFile so the handle is closed even if
reading throws. (B.hGetContents closes the handle on successful
end-of-file, but an exception mid-read would previously leak it.)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 17:26:54 +00:00
John MacFarlaneandClaude ad7b7f9fd9 Text.Pandoc.UTF8: avoid copying input when it contains no CRs.
toText and toTextLazy unconditionally ran a CR-removing filter over
the input, allocating a full copy of the document even in the common
case where no CRs are present. Check for a CR first (B.elem, a fast
memchr) and reuse the input buffer unchanged if none is found; for
the lazy variant, do this chunk-wise to preserve laziness.

On a 10 MB LF-only input this makes toText over 4x faster; when CRs
are present the extra scan is not measurable.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 16:43:52 +00:00
zenor0 2666fbab72 docs: fix fill syntax in Typst property examples (#11855) 2026-09-08 22:44:58 -07:00
John MacFarlaneandClaude 0e399cc799 Update test for raw_html-aware inline <style> handling.
With raw_html disabled (the default for the HTML reader), an inline
<style> element is now ignored rather than passed through as a
RawInline; the raw_html-enabled behavior is still tested.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:12:53 +00:00
John MacFarlaneandClaude 94caebb0a9 HTML reader: fix pre/code attribute precedence for first-wins dedup.
pCodeBlock put the code element's attributes first when merging,
relying on the old last-duplicate-wins behavior of toStringAttr to
give the pre element's attributes precedence.  Now that the first
duplicate wins, put pre's attributes first.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:10:04 +00:00
John MacFarlaneandClaude c1f46cd891 HTML reader: pass the tag name to pSpanLike.
The inline dispatcher already knows which span-like element it is
looking at, so there is no need for pSpanLike to try a parser for
every element in htmlSpanLikeElements.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:08:25 +00:00
John MacFarlaneandClaude ba416103e7 htmlTag: don't copy the remaining input on each invocation.
A space was appended to the remaining input to guarantee a
TagPosition token after the parsed tag; since the input is a strict
Text, this copied the entire remaining input every time htmlTag was
called (e.g. for every inline HTML tag in a markdown document),
giving quadratic behavior in tag-dense documents.  Instead, handle
the case where the tag is the final token by computing the
end-of-input position directly.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:07:14 +00:00
John MacFarlaneandClaude 26e3404455 HTML reader: limit iframe nesting depth.
pIframe fetches the iframe's src and parses it recursively; a
cycle of iframes embedding each other caused infinite recursion.
Track the nesting depth and skip iframes (with a warning) more
than 5 levels deep.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:05:33 +00:00
John MacFarlaneandClaude 4be508863c HTML reader: find list style anywhere in the class attribute.
Previously `<ol class="fancy lower-roman">` got DefaultStyle,
because the whole class attribute was compared against the known
style names.  Check each class individually.  As a side effect, an
unrecognized class no longer prevents falling back to the style
attribute.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:03:51 +00:00
John MacFarlaneandClaude 729a66fa9a HTML reader: report the noteref position for unresolved notes.
The ReferenceNotFound warning was logged after the whole document
had been parsed, so it always pointed at the end of the input.
Record the position of the first reference to each note and use it
in the warning.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:03:07 +00:00
John MacFarlaneandClaude 184e114982 HTML reader: use a Map for the note table.
Avoids a linear scan per noteref (quadratic overall in the number
of notes).  As before, the most recently parsed note with a given
identifier wins.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:01:37 +00:00
John MacFarlaneandClaude ef16134d91 HTML reader: respect raw_html for inline <style> elements.
Inline <style> elements (invalid but common, see #10643) were
always turned into RawInline, even with the raw_html extension
disabled.  Skip them with a warning in that case, as is already
done for block-level style elements.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 03:00:18 +00:00
John MacFarlaneandClaude 910130eff3 HTML reader: let the first of duplicate attributes win.
toStringAttr deduplicated attributes with a right fold, so the last
occurrence of a duplicated attribute was kept.  HTML specifies that
the first occurrence wins (and that lang takes precedence over
xml:lang).  Deduplicate left to right with a seen-set instead.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-09 02:59:30 +00:00