PDF Library for Delphi removes an identity text matrix operator, 1 0 0 1 0 0 Tm, during its save-time peephole content-stream optimization only when the text matrix and text line matrix are already the identity: right after BT, or right after an earlier identity Tm. An identity cm is still always dropped, because cm multiplies the CTM while Tm replaces both text matrices outright. Since v3.539.28 every other identity Tm stays in the stream
The bug this fixes is the quiet kind. A report generator emits BT (Invoice) Tj 1 0 0 1 0 0 Tm (Total) Tj ET, relying on the identity Tm to send the second string back to the text-space origin before it applies its own positioning logic. The older optimizer saw six numbers that spell the identity matrix, decided the operator could not possibly change anything, and deleted it. Nothing failed, nothing logged a warning, and the saved page drew "Total" immediately after "Invoice" on the same baseline, which is exactly the class of defect nobody notices until a customer prints the PDF
Why is 1 0 0 1 0 0 Tm not always a no-op?
An identity Tm is a no-op only when it would replace two matrices that already hold the identity, and that is a property of the operators before it, not of its own operands. ISO 32000-1 §9.4.1 says BT initializes both the text matrix (Tm) and the text line matrix (Tlm) to the identity, and §9.4.2 defines Tm as setting both of them to the given values, not concatenating onto them. Compare that with cm (§8.4.4), which right-multiplies the current transformation matrix: multiplying by the identity leaves any CTM unchanged, so 1 0 0 1 0 0 cm is safe to delete anywhere. Inside a text object the picture is different. Td, TD, T* and a non-identity Tm all move Tlm, and every text-showing operator (Tj, TJ, ', ") advances Tm by the width of the glyphs it painted. After any of those, an identity Tm is a genuine reset to the origin. If you have ever traced text positions by hand with the content-stream CTM and text matrix state tracker, this is the same distinction between concatenating state and replacing it
How the backward scan decides which identity Tm to drop
TPDFContentPeepholeOptimizer.RemoveIdentityMatrices now walks backward from each identity Tm and drops it only if the scan reaches BT or another identity Tm first. The earlier identity Tm counts whether it is kept or was itself just scheduled for deletion, because either way it left both matrices at the identity, exactly as BT does. The rule sorts every operator it can meet into one of two groups:
- Stop and keep the Tm:
Td,TD,T*, a non-identityTm,Tj,TJ,',",ET, any operator the parser does not recognize, or the start of the stream - Step over and keep scanning: operators that never touch Tm or Tlm, such as
Tf,Tc, color setters,gs, marked-content operators andcm
The conservative cases are deliberate. An unknown operator could be anything, so the scan refuses to reason past it. ET closes the text object, so a Tm after it has no BT vouching for the matrix values. The scan also works on one content stream at a time, which matters for pages whose /Contents is an array: a layer that starts in the middle of a text object, with no BT of its own, keeps its identity Tm even when the previous layer would have made it redundant. That costs a few bytes on odd files and never moves a glyph. If you edit page text at the instruction level, as in the character-to-content-byte mapping walkthrough, the same parsed TPDFContentProgram model is what the optimizer rewrites
uses
PDFlibContentModel, PDFlibContentOptimize;
function OptimizeSnippet(const Source: AnsiString): AnsiString;
var
Prog: TPDFContentProgram;
Optimizer: TPDFContentPeepholeOptimizer;
begin
Result := Source;
Prog := TPDFContentProgram.Create;
try
if not Prog.Parse(Source) then
Exit; // damaged stream: leave the bytes alone
Optimizer := TPDFContentPeepholeOptimizer.Create(Prog);
try
Optimizer.Run; // returns the number of instructions removed
finally
Optimizer.Free;
end;
Result := Prog.Emit; // one instruction per line
finally
Prog.Free;
end;
end;
// Removed: Tm directly after BT, the second of two identity Tm in a row
// OptimizeSnippet('BT /F1 12 Tf 1 0 0 1 0 0 Tm (hello) Tj ET')
// Kept: Tm after Td, after Tj, after a non-identity Tm, or outside BT
// OptimizeSnippet('BT (Invoice) Tj 1 0 0 1 0 0 Tm (Total) Tj ET')
Run the helper on the invoice stream from the opening and the identity Tm survives, because the backward scan hits Tj before it reaches BT. Put /F1 12 Tf, 2 Tc and 0 g between BT and the identity Tm and it still goes, since none of those touch the text matrices. A sequence such as BT 10 20 Td 1 0 0 1 0 0 Tm 1 0 0 1 0 0 Tm loses exactly one operator: the first identity Tm resets the matrix that Td moved, and only the second one is redundant
When does the peephole optimizer actually run?
The optimizer runs only during the compression pass, inside TPDFPageTree.Compress, and only on content streams that are not already Flate-compressed. TPDFlib.SetOptimizeContentStreams(1) is the default, and the same switch is exposed as the OptimizeContentStreams field of TPDFlibSaveOptions; both CompressContent and CompressPage honor it. A stream whose /Filter is already /FlateDecode is skipped entirely, so loading an existing compressed PDF and saving it again does not rewrite its operators. If the stream fails to parse, the original decoded bytes are compressed unchanged. TPDFlib.NormalizeContentStreams parses and re-emits content with canonical spacing and numbers but never calls the optimizer, which makes it a useful baseline when you want to see how much of a size difference the peephole rules contribute, alongside the larger wins covered in PDF file size optimization with font subsetting
var
Lib: TPDFlib;
Options: TPDFlibSaveOptions;
begin
Lib := TPDFlib.Create;
try
if Lib.LoadFromFile('report.pdf', '') <> 1 then
Exit;
// Uncompressed streams go through the peephole rules, then Flate
Lib.SetOptimizeContentStreams(1);
Lib.CompressContent;
Lib.SaveToFile('report-optimized.pdf');
// The same choice through the bundled save options; False opts out
Options.CompressContent := True;
Options.CompressFonts := True;
Options.CompressImages := True;
Options.Linearize := False;
Options.KeepModDate := False;
Options.OptimizeContentStreams := False;
Options.GarbageCollect := False;
Options.PackObjectStreams := True;
Lib.SaveToFileOptions('report-plain.pdf', Options);
finally
Lib.Free;
end;
end;
What did the old regression test really guarantee?
The old regression test guaranteed one shape only: an identity Tm directly after BT gets removed. Peephole_RemovesIdentityTextMatrix feeds BT 1 0 0 1 0 0 Tm (hello) Tj ET to the optimizer and asserts that no Tm remains. An earlier release had already noted that dropping an identity Tm is unsafe when Tlm is not the identity, then kept the behavior anyway because the test "locked" it. Reread with care, the test says nothing about an identity Tm after Td or after shown text; treating the coverage of one sample as the contract of the whole rule was the actual mistake. The fix keeps that original case passing and adds six cases that pin both the removable shapes and the kept ones, including a Tm outside any text object and one that follows ET
The trade-off is easy to accept once it is written down. Generators that wrap every text object as BT 1 0 0 1 0 0 Tm ... still get that redundant operator removed, which is where nearly all of the savings came from. What the optimizer gives up is the occasional identity Tm in the middle of a text object, a handful of bytes per page before Flate even sees them, in exchange for a guarantee the module header states plainly: every transform is output-equivalent and never changes the visible page. A size optimizer that moves text is not an optimizer, it is a rendering bug with good compression ratios
The content-stream parser, the peephole optimizer and the save-time compression options described here all ship in the PDF Library for Delphi and C++Builder, which also exposes NormalizeContentStreams, CompressContent and TPDFlibSaveOptions for tuning how each document is written