Technical Article

Identity Tm in PDF Content Streams: Safe Peephole Removal

PDF Library for Delphi removes an identity text matrix operator, 1 0 0 1 0 0 Tm, during its save-time peephole content-stream optimisation only when the text matrix and text line matrix are already the identity: right after BT, or right after an earlier identity Tm. An identity cm is still always dropped, because cm multiplies the CTM while Tm replaces both text matrices outright. Since v3.539.28 every other identity Tm stays in the stream

The bug this fixes is the quiet kind. A report generator emits BT (Invoice) Tj 1 0 0 1 0 0 Tm (Total) Tj ET, relying on the identity Tm to send the second string back to the text-space origin before it applies its own positioning logic. The older optimiser saw six numbers that spell the identity matrix, decided the operator could not possibly change anything, and deleted it. Nothing failed, nothing logged a warning, and the saved page drew "Total" immediately after "Invoice" on the same baseline, which is exactly the class of defect nobody notices until a customer prints the PDF

Why is 1 0 0 1 0 0 Tm not always a no-op?

An identity Tm is a no-op only when it would replace two matrices that already hold the identity, and that is a property of the operators before it, not of its own operands. ISO 32000-1 §9.4.1 says BT initialises both the text matrix (Tm) and the text line matrix (Tlm) to the identity, and §9.4.2 defines Tm as setting both of them to the given values, not concatenating onto them. Compare that with cm (§8.4.4), which right-multiplies the current transformation matrix: multiplying by the identity leaves any CTM unchanged, so 1 0 0 1 0 0 cm is safe to delete anywhere. Inside a text object the picture is different. Td, TD, T* and a non-identity Tm all move Tlm, and every text-showing operator (Tj, TJ, ', ") advances Tm by the width of the glyphs it painted. After any of those, an identity Tm is a genuine reset to the origin. If you have ever traced text positions by hand with the content-stream CTM and text matrix state tracker, this is the same distinction between concatenating state and replacing it

PDFlibPas treats 1 0 0 1 0 0 cm and 1 0 0 1 0 0 Tm differently: cm right-multiplies the CTM and is a no-op anywhere, while Tm replaces Tm and Tlm outright, and every Tj advances Tm by the width it painted, so an identity Tm after shown text is a genuine reset
The report generator counted on that reset: deleting the identity Tm painted Total immediately after Invoice on the same baseline, and nothing failed, logged or warned on the way to the customer's printer

How the backward scan decides which identity Tm to drop

TPDFContentPeepholeOptimizer.RemoveIdentityMatrices now walks backwards from each identity Tm and drops it only if the scan reaches BT or another identity Tm first. The earlier identity Tm counts whether it is kept or was itself just scheduled for deletion, because either way it left both matrices at the identity, exactly as BT does. The rule sorts every operator it can meet into one of two groups:

  • Stop and keep the Tm: Td, TD, T*, a non-identity Tm, Tj, TJ, ', ", ET, any operator the parser does not recognise, or the start of the stream
  • Step over and keep scanning: operators that never touch Tm or Tlm, such as Tf, Tc, colour setters, gs, marked-content operators and cm
PDFlibPas RemoveIdentityMatrices walks backward from each identity Tm: Tf, Tc, colour setters, gs and cm are stepped over, while Td, TD, T*, a non-identity Tm, Tj, TJ, an unknown operator or ET stops the scan and keeps the Tm, and BT vouches for dropping it
An earlier identity Tm stops the scan too, because kept or already scheduled for deletion it left both matrices at the identity — either way the optimiser never moves a glyph

The conservative cases are deliberate. An unknown operator could be anything, so the scan refuses to reason past it. ET closes the text object, so a Tm after it has no BT vouching for the matrix values. The scan also works on one content stream at a time, which matters for pages whose /Contents is an array: a layer that starts in the middle of a text object, with no BT of its own, keeps its identity Tm even when the previous layer would have made it redundant. That costs a few bytes on odd files and never moves a glyph. If you edit page text at the instruction level, as in the character-to-content-byte mapping walkthrough, the same parsed TPDFContentProgram model is what the optimiser rewrites

uses
  PDFlibContentModel, PDFlibContentOptimize;

function OptimizeSnippet(const Source: AnsiString): AnsiString;
var
  Prog: TPDFContentProgram;
  Optimizer: TPDFContentPeepholeOptimizer;
begin
  Result := Source;
  Prog := TPDFContentProgram.Create;
  try
    if not Prog.Parse(Source) then
      Exit; // damaged stream: leave the bytes alone
    Optimizer := TPDFContentPeepholeOptimizer.Create(Prog);
    try
      Optimizer.Run; // returns the number of instructions removed
    finally
      Optimizer.Free;
    end;
    Result := Prog.Emit; // one instruction per line
  finally
    Prog.Free;
  end;
end;

// Removed: Tm directly after BT, the second of two identity Tm in a row
//   OptimizeSnippet('BT /F1 12 Tf 1 0 0 1 0 0 Tm (hello) Tj ET')
// Kept: Tm after Td, after Tj, after a non-identity Tm, or outside BT
//   OptimizeSnippet('BT (Invoice) Tj 1 0 0 1 0 0 Tm (Total) Tj ET')

Run the helper on the invoice stream from the opening and the identity Tm survives, because the backward scan hits Tj before it reaches BT. Put /F1 12 Tf, 2 Tc and 0 g between BT and the identity Tm and it still goes, since none of those touch the text matrices. A sequence such as BT 10 20 Td 1 0 0 1 0 0 Tm 1 0 0 1 0 0 Tm loses exactly one operator: the first identity Tm resets the matrix that Td moved, and only the second one is redundant

When does the peephole optimiser actually run?

The optimiser runs only during the compression pass, inside TPDFPageTree.Compress, and only on content streams that are not already Flate-compressed. TPDFlib.SetOptimizeContentStreams(1) is the default, and the same switch is exposed as the OptimizeContentStreams field of TPDFlibSaveOptions; both CompressContent and CompressPage honour it. A stream whose /Filter is already /FlateDecode is skipped entirely, so loading an existing compressed PDF and saving it again does not rewrite its operators. If the stream fails to parse, the original decoded bytes are compressed unchanged. TPDFlib.NormalizeContentStreams parses and re-emits content with canonical spacing and numbers but never calls the optimiser, which makes it a useful baseline when you want to see how much of a size difference the peephole rules contribute, alongside the larger wins covered in PDF file size optimisation with font subsetting

PDFlibPas runs the peephole optimiser only inside the save-time compression pass: TPDFPageTree.Compress honours SetOptimizeContentStreams, a stream already filtered with /FlateDecode is skipped entirely, an unparseable stream is compressed with its original bytes unchanged, and NormalizeContentStreams never calls the optimiser at all
Skipped compressed streams are the quiet part: load an existing PDF, save it again, and its operators come out untouched because the optimiser only rewrites streams it decoded first
var
  Lib: TPDFlib;
  Options: TPDFlibSaveOptions;
begin
  Lib := TPDFlib.Create;
  try
    if Lib.LoadFromFile('report.pdf', '') <> 1 then
      Exit;
    // Uncompressed streams go through the peephole rules, then Flate
    Lib.SetOptimizeContentStreams(1);
    Lib.CompressContent;
    Lib.SaveToFile('report-optimized.pdf');

    // The same choice through the bundled save options; False opts out
    Options.CompressContent := True;
    Options.CompressFonts := True;
    Options.CompressImages := True;
    Options.Linearize := False;
    Options.KeepModDate := False;
    Options.OptimizeContentStreams := False;
    Options.GarbageCollect := False;
    Options.PackObjectStreams := True;
    Lib.SaveToFileOptions('report-plain.pdf', Options);
  finally
    Lib.Free;
  end;
end;

What did the old regression test really guarantee?

The old regression test guaranteed one shape only: an identity Tm directly after BT gets removed. Peephole_RemovesIdentityTextMatrix feeds BT 1 0 0 1 0 0 Tm (hello) Tj ET to the optimiser and asserts that no Tm remains. An earlier release had already noted that dropping an identity Tm is unsafe when Tlm is not the identity, then kept the behaviour anyway because the test "locked" it. Reread with care, the test says nothing about an identity Tm after Td or after shown text; treating the coverage of one sample as the contract of the whole rule was the actual mistake. The fix keeps that original case passing and adds six cases that pin both the removable shapes and the kept ones, including a Tm outside any text object and one that follows ET

The trade-off is easy to accept once it is written down. Generators that wrap every text object as BT 1 0 0 1 0 0 Tm ... still get that redundant operator removed, which is where nearly all of the savings came from. What the optimiser gives up is the occasional identity Tm in the middle of a text object, a handful of bytes per page before Flate even sees them, in exchange for a guarantee the module header states plainly: every transform is output-equivalent and never changes the visible page. A size optimiser that moves text is not an optimiser, it is a rendering bug with good compression ratios

The content-stream parser, the peephole optimiser and the save-time compression options described here all ship in the PDF Library for Delphi and C++Builder, which also exposes NormalizeContentStreams, CompressContent and TPDFlibSaveOptions for tuning how each document is written