Pod ("Plain old documentation") is a markup languaged used principally
to document Perl modules and programs. Since it was originally meant to
be translated pretty directly to man, the semantics are fairly simple.
This Pod reader was developed with reference to the canonical user and
implementer documentation of Pod: https://perldoc.perl.org/perlpod and
https://perldoc.perl.org/perlpodspec.
There are 1490 .pod, .pl, and .pm in the Perl 5.34 distribution found in
/System/Library/Perl on my mac. Of those, this reader dies with a parse
error on 7 of them. All of them seem to be cases where pod commands are
found within a non-colon-prefixed =begin/=end. perlpodspec says I may
treat this as an error.
[API change] adds readPod
This can be overridden by a final sectPr element in the body
of the reference.docx.
It will only change things for `--top-level-division=chapter`,
since only top-level chapters are put in separate sections.
For that use it will mean that footnote numbers start over with
each chapter, which is usually what is wanted.
Closes#2773.
When `--top-level-division=chapter` is used, a paragraph with
section properties is inserted before each level-1 heading.
By default, this causes the new heading to start on a new page
(though this default can be adjusted in Word).
This change should also make it possible to number footnotes
by chapter (#2773), though that change isn't yet made.
...in the same way it works for other formats (with the top-level
heading being promoted to metadata title). This needed special
treatment because of the way djot surrounds sections with Divs.
Closes#10459.
Avoid the slow URI parser from network-uri on large data URIs.
See #10075. In a benchmark with a large base64 image in HTML ->
docx, this patch causes us to go from 7942 GCs to 3654, and from
3781M in use to 1396M in use.
(Note that before the last few commits, this was running 9099 GCs
and 4350M in use.)
This avoids calling the slow URI parser from network-uri on
data URIs, instead calling our own parser.
Benchmarks on an html -> docx conversion with large base64 image:
GCs from 7942 to 6695, memory in use from 3781M to 2351M,
GC time from 7.5 to 5.6.
See #10075.
The intent is to allow bash process substitution: e.g.,
`--template <(echo "foo")`.
Previously pandoc *always* added an extension based on the
output format, which caused problems with the absolute filenames
used by bash process substitution (e.g. `/dev/fd/11`).
Now, if the template has no extension, pandoc will first
try to find it without the extension, and then add the
extension if it can't be found.
So, in general, extensionless templates can now be used.
But this has been implemented in a way that should not cause
problems for existing uses, unless you are using a template
`NAME.FORMAT` but happen to have an extensionless file `NAME` in
the template search path.
Closes#5270.
This was done to determine the "media category," but we can
get that directly from the mime component of data: URIs.
Profiling revealed that a significant amount of time was
spent in this function when a file contained images with
large data URIs.
Contributes to addressing #10075.
Text.Pandoc.URI: export `pBase64DataURI`. Modify `isURI` to use this
and avoid calling network-uri's inefficient `parseURI` for data URIs.
Markdown reader: use T.P.URI's `pBase64DataURI` in parsing data
URIs.
Partially addresses #10075.
Obsoletes #10434 (borrowing most of its ideas).
Co-authored-by: Evan Silberman <evan@jklol.net>
This patch borrows some code from @silby's PR #10434 and should
be regarded as co-authored. This is a lighter-weight patch
that only touches the Markdown reader.
The basic idea is to speed up parsing of base64 URIs by parsing
them with a special path. This should improve the problem
noted at #10075.
Benchmarks (optimized compilation):
Converting the large test.md from #10075 (7.6Mb embedded image)
from markdown to json,
before: 6182 GCs, 1578M in use, 5.471 MUT, 1.473 GC
after: 951 GCs, 80M in use, .247 MUT, 0.035 GC
For now we leave #10075 open to investigate improvements in
HTML rendering with these large data URIs.
Co-authored-by: Evan Silberman <evan@jklol.net>
We used to set this to `.html`, but this seemed inappropriate
once we started using this function for `--pdf-engine=typst`.
So we changed it in pandoc 3.6 to `.source`. But apparently
`wkhtmltopdf` needs it to be `.html`. So now we have added
a parameter to `toPdfViaTempFile` that allows the extension
to be specified.
Closes#10468.