PDF Library for Delphi (PDFlibPas) versions before v3.539.47 could decode escaped text twice when drawing HTML or Markdown into a PDF. DrawHTMLText and DrawHTMLTextBox parse the HTML, normalize it back into HTML, then parse it again, so text written as <unsafe> reached the second parse as a real tag. Since v3.539.47 each entity is decoded exactly once and text is re-escaped wherever it turns back into HTML
The scenario that exposes this is ordinary. A help desk exports tickets to PDF, and the customer comment goes into an HTML template. The developer did the right thing and escaped the comment, so <b> became <b>. Inside the renderer that escaping was quietly undone: the comment came out bold, an unknown tag name simply vanished from the page, and an escaped anchor turned into a clickable link annotation. No exception, no warning, a perfectly valid PDF that says something different from the data
Why does escaped text become a real tag in the PDF?
Escaped text became markup because the renderer runs two parse passes, and the normalization step between them wrote already decoded text back into HTML without escaping it again. Every decode that the first parse performed was then available to the second parse as live syntax
The two passes exist for a good reason. The first parse builds a list of tag and word elements. NormalizeParsedHTML then resolves the stylesheet cascade: it matches the rules from <style> blocks against each tag, merges them with inline style attributes, stores the result on the tag, and serializes the whole element list back into an HTML string. The layout pass parses that normalized string. It is the same machinery that drives flexbox, CSS grid and footnote layout in PDFlibPas HTML rendering
The flaw was in how words were serialized. Tags were written back from their original source form, while words were written back in their decoded form. A word that the first parse had decoded from <unsafe> to <unsafe> landed in the normalized HTML as raw angle brackets, and the second parse read it as an element. Around that core bug sat three smaller leaks that pointed the same way:
&was not in the supported entity set, soR&Dprinted literally and there was no way to write a literal entity spelling such as<as text- The drawing stage replaced
a second time, after parsing had already finished, so a literal entity spelling could still disappear at the very end - Markdown code escaping skipped the ampersand, and the dataset exporter escaped only angle brackets, so entity spellings inside code or cell values were decoded as markup
| Input reaching the renderer | Before v3.539.47 | Since v3.539.47 |
|---|---|---|
<unsafe> | Parsed as a tag, text never reaches the page | <unsafe> drawn as text |
<b>x</b> | x drawn in bold | <b>x</b> drawn as text |
R&D | R&D printed literally | R&D |
&lt; | &lt; printed literally | < |
Markdown code span containing | Became a non-breaking space | drawn as text |
Dataset cell value < | < | < |
How v3.539.47 makes HTML entity decoding single-pass
PDFlibPas v3.539.47 makes entity decoding single-pass with three coordinated changes: the parser decodes & last, the drawing stage no longer decodes anything, and every place that turns decoded words back into HTML escapes them again first
The supported entity set for text content is now <, >, & and . Anything else, including numeric references such as A and named entities such as ", stays literal text. That boundary matters for how you escape your own input, as shown below
Order inside the decoder is the first fix. If & were decoded first, the input &lt; would become < and the next replacement would turn it into <, a double decode that happens inside one pass. The ANSI word path therefore replaces <, > and first and & last, so the ampersand it produces is never examined again. The UTF-16 word path is a single left-to-right scan in two-byte steps that rewrites each match in place and moves past it, which gives the same guarantee structurally
The second fix removes the late replacement from the drawing stage. Decoding belongs to the parser and nowhere else, so a word that reaches the line breaker is final text
The third fix is the boundary rule. NormalizeParsedHTML now escapes &, < and > in every decoded word before appending it to the normalized HTML. The second parse decodes it back to exactly the same text, so the net effect over the whole pipeline is one decode. The continuation string follows the same rule: words that did not fit into the box are escaped before they are appended to LeftOverText, and the rest of the remainder is copied from the normalized HTML, which is already in escaped form. The loop that collects those leftover words is also bounded by the word count now, where the old repeat loop could step past the last word
Why can't UTF-16BE escaping use a byte-level replace?
UTF-16BE escaping cannot use a byte-level replace because the two-byte pattern for an ampersand can straddle two unrelated characters. The only correct unit of work is the whole 16-bit code unit
The renderer stores Unicode words as big-endian UTF-16 packed into byte strings, high byte first. An ampersand is 00 26. Now take U+0100 (Latin capital A with macron, bytes 01 00) followed by U+2603 (the snowman, bytes 26 03). The byte sequence is 01 00 26 03, and bytes two and three read 00 26. A byte search for #0'&' finds an ampersand that does not exist, splices the bytes for & into the middle of two characters, and shears every following character by one byte
That is not an exotic corner case. Any character whose low byte is zero can supply the first half; U+4E00, one of the most frequent CJK ideographs, qualifies. The angle brackets have the same exposure: 00 3C and 00 3E appear whenever such a character is followed by one from U+3C00 to U+3EFF in CJK Extension A. The fix in EscapeHTMLWord unpacks the bytes to a WideString, escapes character by character and packs the result again. The decoder side was already safe because it only tests patterns at even code unit boundaries
The same rule applies to your own code. If you ever hold UTF-16 text as TBytes, for example after TEncoding.BigEndianUnicode.GetBytes, do not search it for byte patterns. Convert back to a string and work on characters
Markdown code blocks and dataset exports: escape the ampersand first
Since v3.539.47 both HTML producers inside PDFlibPas, the Markdown converter and the dataset exporter, escape the ampersand before the angle brackets, so the single decode in the renderer restores exactly the original text
In MarkdownToHTML, inline code spans and fenced or indented code blocks now map & to &, < to < and > to >, while spaces become and a tab becomes four of them to keep indentation. Ordinary Markdown prose escapes only the angle brackets, so raw HTML in prose cannot inject tags while an author can still write & on purpose, much as Markdown authors expect. DrawMarkdownText and DrawMarkdownTextBox use the same conversion, so code shows up in the PDF exactly as typed:
uses
System.SysUtils, PDFlibrary;
procedure RenderCodeSample;
var
Lib: TPDFlib;
Md, Html: WideString;
begin
Md := 'Comparison helper:' + sLineBreak + sLineBreak +
'```' + sLineBreak +
'if (A < B) and (Flags <> 0) then' + sLineBreak +
' WriteLn(''<tag> & R&D'');' + sLineBreak +
'```';
Lib := TPDFlib.Create;
try
// Inspect the HTML: in code, '&' becomes '&' and '<' becomes '<'
Html := Lib.MarkdownToHTML(Md);
Lib.SetOrigin(1); // top-left origin, Y grows downward
Lib.SetMeasurementUnits(0); // points
// The page shows the code exactly as typed, entity spellings included
Lib.DrawMarkdownText(50, 50, 495, Md);
Lib.SaveToFile('code-sample.pdf');
finally
Lib.Free;
end;
end;
The dataset exporter is the instructive case. Before v3.539.47 it escaped only angle brackets, and on purpose: the renderer did not decode &, so escaping the ampersand would have printed & in every cell containing one. The workaround was correct for the old renderer and wrong in general, because a cell value that happened to contain < was decoded into <. With the renderer fixed, the exporter escapes & first, and a value such as R&D < & lands in the PDF verbatim. If you build reports that way, the walkthrough on exporting a TDataSet to a PDF report in Delphi covers the rest of the exporter
Why the ampersand must go first is worth spelling out once. Escape < first and you get <; escape & second and that becomes &lt;, which a correct single decode displays as < instead of <. A sequential replace chain is only correct when the escape character itself is handled before anything that introduces it
How should you escape untrusted text for DrawHTMLTextBox?
For PDFlibPas HTML rendering, escape untrusted text content by replacing &, then <, then >, exactly once, and keep untrusted data out of attribute values entirely
uses
System.SysUtils, PDFlibrary;
// Escapes untrusted text for PDFlibPas HTML text content.
// '&' must be replaced first, otherwise the ampersand inside
// an already produced '<' would be escaped a second time
function EscapeHTMLText(const S: string): string;
begin
Result := StringReplace(S, '&', '&', [rfReplaceAll]);
Result := StringReplace(Result, '<', '<', [rfReplaceAll]);
Result := StringReplace(Result, '>', '>', [rfReplaceAll]);
end;
procedure RenderTicket(const CustomerComment: string);
var
Lib: TPDFlib;
Html: WideString;
begin
Lib := TPDFlib.Create;
try
Lib.SetOrigin(1);
Lib.SetMeasurementUnits(0);
Html := '<p><b>Customer comment</b></p>' +
'<p>' + EscapeHTMLText(CustomerComment) + '</p>';
Lib.DrawHTMLText(50, 50, 495, Html);
Lib.SaveToFile('ticket.pdf');
finally
Lib.Free;
end;
end;
On v3.539.47 a comment such as Try <a href="https://example.com">this</a> & <b> appears on the page character for character. Before v3.539.47 the same escaped input could produce a live link annotation, which is the part that turns a display glitch into a security problem: a ticket comment should never be able to plant a clickable URL in a document your staff trusts
Notice what the function does not escape. General-purpose HTML escapers also convert " to " and ' to ', which is right for a browser. PDFlibPas text decoding recognizes only the four entities listed earlier, so those two would print literally as " and '. Quotes are harmless in text content; they only matter inside attribute values, and the renderer does not decode entities in attributes at all. The safe design is therefore not a better escaper but a rule: untrusted data never goes into href, src or style. If a link target really must come from user data, validate it yourself against an allow-list of schemes and characters and reject anything containing quotes or angle brackets
Two upgrade notes follow directly from the fix:
- If your code stopped escaping
&because older versions printed&literally, add it back. Without it, user text containing<now displays as<, still harmless text but no longer what the user typed - Do not escape twice. Text that passes through two escapers renders
<as the visible spelling<, so find the one boundary where your data enters HTML and escape there only
Paginating with LeftOverText without breaking escapes
DrawHTMLTextBox returns the HTML that did not fit, usually called LeftOverText, and since v3.539.47 that remainder preserves literal entity spellings and escaped angle brackets when you pass it to the next box. The rule for callers is simple: pass it back unchanged
const
BoxLeft = 50;
BoxTop = 50;
BoxWidth = 495; // sized for an A4 page in points
BoxHeight = 740;
MaxPages = 500;
procedure RenderLongHTML(Lib: TPDFlib; const Html: WideString);
var
Rest: WideString;
Pages: Integer;
begin
Lib.SetOrigin(1);
Lib.SetMeasurementUnits(0);
Rest := Lib.DrawHTMLTextBox(BoxLeft, BoxTop, BoxWidth, BoxHeight, Html);
Pages := 1;
while (Rest <> '') and (Pages < MaxPages) do
begin
Lib.NewPage;
Inc(Pages);
// LeftOverText is already escaped engine HTML: never escape or unescape it
Rest := Lib.DrawHTMLTextBox(BoxLeft, BoxTop, BoxWidth, BoxHeight, Rest);
end;
if Rest <> '' then
raise Exception.CreateFmt('Content still left after %d pages', [MaxPages]);
end;
Treat the remainder as opaque. It is the normalized HTML of the engine, with styles already resolved, so do not run it through your own escaper, do not decode it, and do not splice user text into it. The page cap is cheap insurance: if some element can never fit the box, an uncapped loop has no natural exit
Markdown has its own continuation. DrawMarkdownTextBox returns a token that starts with an internal marker so that the next call can skip conversion; hand it back to DrawMarkdownTextBox or DrawMarkdownText, not to the HTML entry points, which would draw the marker as text
The general lesson: decode once, re-encode at every boundary
Any pipeline that parses text, serializes the result back into the same syntax and parses it again must treat decoding as an operation that happens in exactly one place, and must re-encode at every boundary where decoded text becomes syntax again. Template engines, HTML sanitizers and Markdown-to-HTML-to-PDF chains share this shape and fail the same way when a serializer forgets that it produces markup
The symptoms are predictable once you know the shape. Too little re-encoding turns data into syntax, which is the injection direction. Too much encoding, or a decoder that runs twice, shows entity spellings to the reader or eats them, which is the display direction. Fixing one direction alone usually breaks the other, which is why the PDFlibPas fix had to add & decoding, reorder it, remove the late decode and add re-escaping in the same release. The same principle runs the other way when PDF content is exported as structured text, as in PDF to Markdown and DOCX semantic export from Delphi, where every literal character must be escaped for the target syntax exactly once
Quick reference checklist
- Upgrade to PDFlibPas v3.539.47 or later if you render HTML or Markdown that contains user data
- Escape text content with
&first, then<and>; do not convert quotes for PDFlibPas text - Escape once, at the single point where data enters the HTML string
- Keep untrusted values out of
href,srcandstyle, or validate them against an allow-list - Expect only
<,>,&and to be decoded in text; other entities stay literal - Pass
LeftOverTextback toDrawHTMLTextBoxunchanged and cap the page loop - Pass Markdown continuation tokens only to
DrawMarkdownTextBoxorDrawMarkdownText - Never search UTF-16 byte buffers for byte patterns; work on whole code units
HTML and Markdown rendering, dataset report export and the rest of the layout engine ship in the native Pascal source of PDF Library for Delphi, for Delphi and Free Pascal. See the PDFlibPas product page for editions, platform support and a trial download