Remove old ad hoc `paraToPlain` at the table cell level.
This should ensure that we don't get tables that mix
Plain and Para. Such tables tend to look funny when
rendered in docx and other formats.
Closes#11864.
Previously a paragraph that was, e.g. center-aligned would be
missing its default Body Text style.
Also ensure that any block sequence (list items, table cells,
block quote, div), First Paragraph is set for the first paragraph
in the item.
Closes#11867.
further tweaks
parseXMLContentsWithEntities previously used xml-conduit's document
parser, which requires a single root element, and worked around this
by re-parsing fragments wrapped in "<wrapper>...</wrapper>" when the
first parse failed with ContentAfterRoot. This was fragile: it broke
on fragments with an XML declaration or DOCTYPE before multiple root
elements, failed on some text-only input, parsed fragments twice, and
gave different results depending on which path was taken.
We now drive xml-conduit's streaming parser directly, folding the
event stream into a Content forest in a single pass, with start/end
tag balance checked during the fold (the stream parser itself does
not validate this). Behavior changes:
- Content fragments with an XML declaration or DOCTYPE followed by
multiple root elements, and text-only or empty input, now parse
instead of erroring.
- Attributes now preserve document order instead of being sorted
alphabetically (the document parser stored them in a Map).
- Errors for unresolved entities and mismatched tags now report
source positions.
This is also slightly faster than the old document-parser path.
parseXMLElementWithEntities is unchanged, still requiring a single
root element. The xml-light library component gains a direct
dependency on conduit (already a transitive dependency via
xml-conduit).
Co-Authored-By: Claude <noreply@anthropic.com>
Previously the prettified CData path rendered the escaped text to a
Builder, materialized it as lazy then strict Text, unpacked it to
String, and re-folded it into a Builder one singleton at a time.
Since escaping neither adds nor removes newlines, we can instead
split the unescaped text on newlines, escape each line, and
interleave the indentation, emitting whole chunks.
This makes showCData and ppcCData unused, so they are removed.
Co-Authored-By: Claude <noreply@anthropic.com>
Previously, if a text contained any escapable character, the
whole text was unpacked to String and rebuilt one Builder
singleton per character. Now we copy whole chunks between
escapable characters with fromText, and avoid the extra
initial T.any pass. About 5x faster on typical text with
occasional escapes; output is unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
escCData stopped processing after the first "]]>" occurrence,
emitting the rest of the text verbatim; a second "]]>" would
prematurely terminate the CDATA section, producing invalid XML.
(The original xml-light code recurses after each occurrence; the
recursion was lost in the Text port.) Rewritten with T.breakOn,
which also avoids building the output one character at a time.
Co-Authored-By: Claude <noreply@anthropic.com>
Table row parsers are speculatively tried at every paragraph line
boundary via the terminator continuation. Each of the three row
variants would fully inline-parse the line before failing for lack of
a cell separator, so every paragraph line was inline-parsed several
times.
Since cell separators (whitespace followed by pipes) cannot span
lines, scanning the raw line for one lets us reject non-table lines
without any inline parsing. The check is conservative and never
rejects a line a row parser could accept.
Together with the previous commit, this makes the reader 4.3x faster
on a 1.6 MB test document and reduces heap allocation by 87%.
Co-Authored-By: Claude <noreply@anthropic.com>
Plain words previously fell through ~20 failing alternatives (emphasis,
tag, and link parsers) before reaching str. Since every other
alternative starts with a non-alphanumeric character and str only
matches alphanumerics, trying str right after whitespace cannot change
parse results but avoids the failed attempts for every word.
On a 1.6 MB test document this makes the reader 2.2x faster and
reduces heap allocation by 61%.
Co-Authored-By: Claude <noreply@anthropic.com>
The friendly mediaPath was derived by percent-unescaping the key, so
distinct keys like "a%20b.png" and "a b.png" produced the same
mediaPath ("a b.png") and silently clobbered each other on
extraction (and inside docx/epub archives). Now the original name
is only kept if the key contains no percent sign, so mediaPath
equals the key and distinct keys yield distinct paths; anything
percent-encoded gets a content-hash name. Hashed names can only
coincide for identical contents, which is harmless.
Co-Authored-By: Claude <noreply@anthropic.com>
normalise does not remove redundant path components, so img/../a.png
and a.png were distinct keys: the same resource could be stored
twice, and a lookup by one spelling missed an insert by the other.
Use makeCanonical (as PandocPure's FileTree already does for its
path-indexed map), which also handles duplicate and trailing
slashes, replacing backslashes with slashes first. As a side
effect, paths with a collapsible ".." now keep a friendly name
instead of being renamed to a content hash.
Co-Authored-By: Claude <noreply@anthropic.com>
The insertMedia check used isInfixOf, so a harmless name like
foo..bar.png was silently renamed to its content hash. Check for an
actual ".." path component instead, treating both / and \ as
separators since the unescaped path may contain backslashes from
percent-encoding. Added tests.
Co-Authored-By: Claude <noreply@anthropic.com>
Network.URI.isURI treats Windows drive-letter paths like c:/foo.png
as URIs (the drive letter parses as a scheme), so they escaped path
normalization; and it rejects URIs containing non-ASCII characters,
so e.g. https://example.com/résumé.png fell into the file-path
branch, where normalise collapsed the double slash. Pandoc's own
isURI restricts to known schemes, escapes non-ASCII characters
before parsing, and has a fast path for base64 data URIs.
Co-Authored-By: Claude <noreply@anthropic.com>
Use strict fields for mediaMimeType and mediaPath, and switch to
Data.Map.Strict so that insertMedia forces them at insert time.
Previously the lazy map stored an unevaluated record construction
that retained the original path, the parseURI result, and the
mime-type fallback machinery until first lookup. mediaContents is
left lazy so contents need not be forced at insert time.
Co-Authored-By: Claude <noreply@anthropic.com>
URI schemes are case-insensitive (RFC 3986, section 3.1), but several
places matched only lowercase "data:" or "file:":
- Text.Pandoc.URI.pBase64DataURI (also used by isURI)
- downloadOrRead: the fast-path prefix check and the parseURI
fallback for data: and file: URIs
- Text.Pandoc.App.Input.readSource: file:, http:, and https: URIs
given as input sources
- Text.Pandoc.MediaBag: canonicalize's fast path and insertMedia's
data-URI branch
Previously, e.g. DATA:image/gif;base64,... would be handed to openURL
and the image dropped, and a large uppercase data URI would also force
a full parseURI in MediaBag.canonicalize. Now such URIs are handled
the same as their lowercase equivalents. Added a test case.
Co-Authored-By: Claude <noreply@anthropic.com>
Hash the lazy bytestring chunk-wise instead of forcing the entire
contents into a single strict buffer with BL.toStrict. Digests are
unchanged.
Co-Authored-By: Claude <noreply@anthropic.com>
Commit 455bea907 added the new module but did not register it in
pandoc.cabal or remove the original definitions, leaving the file
uncompiled and Text.Pandoc.Class.IO with duplicate copies of
openURL and getManager. Add the module to other-modules and make
Text.Pandoc.Class.IO re-export openURL from it, deleting the
duplicated code and now-unneeded imports.
Co-Authored-By: Claude <noreply@anthropic.com>
The recursive directory traversal followed symlinks without any
cycle detection, so a symlink pointing to an ancestor directory
caused non-termination. Track the canonical paths of ancestor
directories and skip a directory that is already on the current
chain. Symlinks to directories outside the chain are still
followed as before.
Co-Authored-By: Claude <noreply@anthropic.com>
The HTTP manager is created lazily with TLS settings based on
stNoCheckCertificate and then cached in CommonState, so changing the
option after the first request had no effect. Discard the cached
manager when the option's value changes.
Co-Authored-By: Claude <noreply@anthropic.com>
The offset in PandocUTF8DecodingError was computed with B.elemIndex,
i.e. the first occurrence of the offending byte *value* anywhere in
the file. If the same byte occurred earlier as part of a valid
multi-byte sequence, the reported position pointed at valid text.
Scan for the first invalid UTF-8 sequence instead and report its
actual position and byte.
Co-Authored-By: Claude <noreply@anthropic.com>
Percent-escapes in a data URI represent raw octets (RFC 2397), but
the data was decoded with unEscapeString (which UTF-8-decodes
consecutive escapes into Chars) and then re-encoded with
UTF8.fromString. This corrupted percent-encoded binary data:
e.g. %89 became the two bytes C2 89. Decode the escapes directly
to bytes instead. Also split off the media type before unescaping,
so an escaped comma cannot shift the separator.
Co-Authored-By: Claude <noreply@anthropic.com>
The base64 indicator in a data URI is the final parameter of the
media type and may follow other parameters, e.g.
data:text/plain;charset=utf-8;base64,... Previously such URIs were
not base64-decoded, because the code expected ";base64" to be the
only parameter. The media type parameters are now also retained in
the returned MIME type, as they already are in pBase64DataURI.
Co-Authored-By: Claude <noreply@anthropic.com>
The check used a string prefix comparison, so any path beginning
with two dots (e.g. `..foo/bar.yaml`) was treated as pointing
outside the working directory, wrongly disabling the user data
directory fallback in checkUserDataDir. Compare the first path
component instead.
Co-Authored-By: Claude <noreply@anthropic.com>
Previously, if the action passed to runSilently threw an error that
was later caught, the verbosity remained pinned at ERROR and all
previously accumulated log messages were lost. Now the original
log and verbosity are restored even when the action fails.
Co-Authored-By: Claude <noreply@anthropic.com>
Use withFile instead of openFile so the handle is closed even if
reading throws. (B.hGetContents closes the handle on successful
end-of-file, but an exception mid-read would previously leak it.)
Co-Authored-By: Claude <noreply@anthropic.com>
toText and toTextLazy unconditionally ran a CR-removing filter over
the input, allocating a full copy of the document even in the common
case where no CRs are present. Check for a CR first (B.elem, a fast
memchr) and reuse the input buffer unchanged if none is found; for
the lazy variant, do this chunk-wise to preserve laziness.
On a 10 MB LF-only input this makes toText over 4x faster; when CRs
are present the extra scan is not measurable.
Co-Authored-By: Claude <noreply@anthropic.com>
With raw_html disabled (the default for the HTML reader), an inline
<style> element is now ignored rather than passed through as a
RawInline; the raw_html-enabled behavior is still tested.
Co-Authored-By: Claude <noreply@anthropic.com>
pCodeBlock put the code element's attributes first when merging,
relying on the old last-duplicate-wins behavior of toStringAttr to
give the pre element's attributes precedence. Now that the first
duplicate wins, put pre's attributes first.
Co-Authored-By: Claude <noreply@anthropic.com>
The inline dispatcher already knows which span-like element it is
looking at, so there is no need for pSpanLike to try a parser for
every element in htmlSpanLikeElements.
Co-Authored-By: Claude <noreply@anthropic.com>
A space was appended to the remaining input to guarantee a
TagPosition token after the parsed tag; since the input is a strict
Text, this copied the entire remaining input every time htmlTag was
called (e.g. for every inline HTML tag in a markdown document),
giving quadratic behavior in tag-dense documents. Instead, handle
the case where the tag is the final token by computing the
end-of-input position directly.
Co-Authored-By: Claude <noreply@anthropic.com>
pIframe fetches the iframe's src and parses it recursively; a
cycle of iframes embedding each other caused infinite recursion.
Track the nesting depth and skip iframes (with a warning) more
than 5 levels deep.
Co-Authored-By: Claude <noreply@anthropic.com>
Previously `<ol class="fancy lower-roman">` got DefaultStyle,
because the whole class attribute was compared against the known
style names. Check each class individually. As a side effect, an
unrecognized class no longer prevents falling back to the style
attribute.
Co-Authored-By: Claude <noreply@anthropic.com>
The ReferenceNotFound warning was logged after the whole document
had been parsed, so it always pointed at the end of the input.
Record the position of the first reference to each note and use it
in the warning.
Co-Authored-By: Claude <noreply@anthropic.com>
Avoids a linear scan per noteref (quadratic overall in the number
of notes). As before, the most recently parsed note with a given
identifier wins.
Co-Authored-By: Claude <noreply@anthropic.com>
Inline <style> elements (invalid but common, see #10643) were
always turned into RawInline, even with the raw_html extension
disabled. Skip them with a warning in that case, as is already
done for block-level style elements.
Co-Authored-By: Claude <noreply@anthropic.com>
toStringAttr deduplicated attributes with a right fold, so the last
occurrence of a duplicated attribute was kept. HTML specifies that
the first occurrence wins (and that lang takes precedence over
xml:lang). Deduplicate left to right with a seen-set instead.
Co-Authored-By: Claude <noreply@anthropic.com>