GHC 9.14 ships base 4.22 and containers 0.8, newer than the bounds of
serialise, cborg, and haddock-library allow. Relax them under
impl(ghc >= 9.14) so other lanes are unaffected, and force -fhttp on
the 9.14 lane: given the choice, the solver prefers disabling the
http flag to violating bounds.
Inline elements with a content body (emph, strong, underline, strike,
sub, super, smallcaps, highlight, link, lower, upper, ...) failed with
a parse error when the body contained a paragraph break or other block
content, since inline handlers parse their bodies with pInlines, which
cannot consume paragraph breaks. Example:
```
#emph[hi
there]
```
Generalize `fixNesting`, which already hoisted single block children out
of inline elements (#11017): the content of an inline element is now
split at paragraph breaks and at block children; inline runs are
wrapped back in the element, and block children are pulled out with
the styling applied to their own contents when they have a body. A
pandoc inline cannot span paragraphs, so this is the closest
structural rendering; the paragraph structure is preserved.
Remove the kludge we used for `#quote`, using `getInlineBody`; it
is no longer needed with this general fix.
Previously if some columns were ColWidthDefault and others ColWidth 0.x,
we would get columns with a specified width of 0 for the default ones.
Instead, treat all columns as default in this case.
Closes#11899.
A link with a non-empty title now gets a `w:tooltip` attribute on its
`w:hyperlink`, which Word shows as a ScreenTip and the docx reader
reads back as the title.
Closes#11869.
Co-Authored-By: Claude <noreply@anthropic.com>
A hyperlink's ScreenTip is stored as `w:tooltip` on `w:hyperlink`,
for external and internal links alike. It now becomes the title of
the Link instead of being discarded.
The test document was made in Word for Microsoft 365 (Version 2609,
Build 16.0.20430.20032) on Windows 11 25H2 (OS Build 26200.9168).
See #11869.
Co-Authored-By: Claude <noreply@anthropic.com>
...when parsing a block quote or list item. Otherwise
it can happen that by the time parseWithString' is called,
the position has already been set to the next file on the
command line.
Closes#11888.
Closes#11888.
Previously this made three full walks of the document: one for block
attributes, one for inline attributes, and one for link targets.
Fold the link-target fixing into the inline attribute walk, and
inline the now single-use walkAttr helper (not exported), reducing
this to two walks.
Ad hoc benchmark (calling ensureValidXmlIdentifiers on a parsed
8000-section document with ~500k inlines, 20 iterations, forced with
deepseq): 3.87s -> 3.17s (~18% faster). Output is byte-identical
for markdown -> html4.
Co-Authored-By: Claude <noreply@anthropic.com>
Each attribute was rendered with `text . T.unpack`, converting the
escaped Text to String and back. Use `literal . fromText` instead
(free for Doc Text, which is what all callers use). This adds a
FromText constraint to htmlAttrs and tagWithAttrs.
Benchmark (30k-cell table with id/class/data attrs, json ->
mediawiki): 2.15s -> 2.00s with ~70-char attribute values; no change
with short values, where attribute text is a small fraction of the
document.
Co-Authored-By: Claude <noreply@anthropic.com>
endsWithPlain recursed into the last item of BulletList and
OrderedList but ignored DefinitionList, so list items ending with a
compact definition list were treated as loose by the RST, Org, and
Haddock writers.
Co-Authored-By: Claude <noreply@anthropic.com>
An item like `- [ ]` with no following text parses to
`Plain [Str "☐"]` (no Space), which toTaskListItem did not match.
So such items were rendered as a literal ☐/☒ character, and a single
empty item kept the whole list from being treated as a task list
(e.g. the HTML writer's task-list class and checkbox rendering).
Co-Authored-By: Claude <noreply@anthropic.com>
Entries whose RowSpan fell to 0 were left in the map, relying on
every consumer to guard against them. Delete them instead, matching
what takePreviousSpansAtColumn already does. No change in behavior
for current consumers (AsciiDoc, ANSI writers).
Co-Authored-By: Claude <noreply@anthropic.com>
walkAttr only rewrote attributes on Header, CodeBlock, Table, and
Div, while fixLinks rewrote every internal link starting with a
non-letter. As a result, ids on Figure, TableHead, TableBody,
TableFoot, Row, and Cell were left unchanged while links to them were
renamed, producing broken internal links in the HTML4/XHTML, EPUB,
DocBook, TEI, ICML, FB2, and ODT writers.
Co-Authored-By: Claude <noreply@anthropic.com>
Every Code inline rebuilt an association list mapping all ~30
highlighting token types to their rStyle elements (30 style-map
lookups plus string conversions), and each highlighted token then did
a linear scan of that list. Compute the map once at the start of
writing (it depends only on the style maps, which are fixed) and store
it in the writer state as a Data.Map.
On an ad hoc benchmark with 100K inline code spans, conversion time
drops from 2.4s to 1.5s.
Co-Authored-By: Claude <noreply@anthropic.com>
convertSpace merged adjacent Str/Space inlines by repeatedly
concatenating onto the accumulated Text, so a paragraph of n words
required O(n^2) copying. Accumulate the chunks and concatenate once
instead.
On an ad hoc benchmark (4 paragraphs of 80K words each), conversion
time drops from 3.2s to 1.1s, and time no longer depends on paragraph
length (16 x 20K words previously took 1.8s, now also 1.2s).
Co-Authored-By: Claude <noreply@anthropic.com>
Numeric.showHex produces the same lowercase hex output as printf "%x"
without the overhead of format string interpretation.
Co-Authored-By: Claude <noreply@anthropic.com>
Previously the code relied on laziness to avoid the extra full render
of the document body used in the trailer ID hash. Make this explicit
by only computing the trailer ID when it will actually be used.
Co-Authored-By: Claude <noreply@anthropic.com>