Technical Article

FDF Annotation Import in Delphi: Fixing the Silent Zero

Before v3.539.30, TPDFlib.ImportAnnotationsFromFDFString in losLab PDF Library returned the number of FDF annotation entries it had parsed while adding none of them to the document: every entry was counted, every entry was dropped. Since v3.539.30 the FDF importer reads keys in any order, parses /Rect correctly and locale-independently, and the matching exporter writes the annotation's real /Rect, so an export, import and second export produce byte-identical FDF. The rest of this note explains how one wrong starting offset produced a perfect silent failure, which three other defects were hiding behind it, and how to check an import yourself instead of trusting the return value

The scenario is ordinary. A reviewer marks up a contract, the comments travel as an FDF file (Acrobat calls it Export Comments), and your Delphi service merges them into a clean copy with ImportAnnotationsFromFDF. The call returns 7, the log says "7 comments imported", the job goes green, and the output PDF has no comments at all. Nothing raised, nothing warned, and the number looked plausible because it was the true count of entries in the file. That is the worst shape a bug can take: a function whose only signal of success is a counter that is computed independently of the work it claims to report

Why did ImportAnnotationsFromFDFString report success but add nothing?

The importer read every /Subtype as an empty string, and the helper that creates the annotation exits early on an empty subtype while the caller increments the result anyway. The key finder returned the position immediately after /Subtype, which is the whitespace before the value. ReadName started at that space and stopped at the first whitespace character, so it stopped before reading anything. AddAnnotationToPage refuses to build an annotation without a subtype, which is the correct defensive choice in isolation, but it was a procedure with no return value, and Inc(Result) sat outside it. Each guard was reasonable on its own; together they converted "nothing worked" into "everything worked". The fix makes ReadName skip whitespace, require the leading / of a PDF name object, and stop at any delimiter, including [, ( and ), so /Subtype/Text and /Subtype /Text both yield Text

PDFlibPas ImportAnnotationsFromFDFString found /Subtype, started ReadName on the whitespace after the key so it returned an empty name, AddAnnotationToPage exited on the missing subtype, and the caller incremented the result anyway, reporting seven imported comments while adding none to the document
Every guard was reasonable in isolation; together they converted nothing worked into everything worked, which is why the return value must never be the only thing an import test checks

The return value deserved care even after that fix. Through v3.539.39, ImportAnnotationsFromFDFString still incremented its result for every well-formed dictionary in the /Annots array, including entries whose 0-based /Page was out of range or whose /Subtype was missing, both of which are skipped. Since PDFlibPas v3.539.40, ImportAnnotationsFromFDFString and ImportAnnotationsFromFDF return the number of annotations actually added, like the XFDF import: the FDF helper AddAnnotationToPage now returns a Boolean and the counter only moves on success. Measuring the document is still the stronger check, because it also holds on older versions, so the sketch below compares AnnotationCount on each page before and after the import

function TotalAnnotations(Lib: TPDFlib): Integer;
var
  Page, Saved: Integer;
begin
  Result := 0;
  Saved := Lib.SelectedPage;
  for Page := 1 to Lib.PageCount do
    if Lib.SelectPage(Page) = 1 then
      Inc(Result, Lib.AnnotationCount);   // per selected page, widgets included
  Lib.SelectPage(Saved);
end;

var
  Lib: TPDFlib;
  Before, Reported, Added: Integer;
begin
  Lib := TPDFlib.Create;
  try
    Lib.LoadFromFile('contract.pdf', '');
    Before := TotalAnnotations(Lib);
    Reported := Lib.ImportAnnotationsFromFDF('review-comments.fdf');
    Added := TotalAnnotations(Lib) - Before;
    if Added <> Reported then   // equal since v3.539.40
      Writeln(Format('Importer reported %d, %d landed on a page', [Reported, Added]));
    Lib.SaveToFile('contract-reviewed.pdf');
  finally
    Lib.Free;
  end;
end;

Three more defects behind the first one

Fixing the subtype alone would have exposed three further bugs in the same function, each of which had been invisible only because no annotation ever reached a page. First, ReadNumber took its position as a value parameter, so reading the four /Rect numbers in sequence read the same spot four times, and it did not skip the opening [, so in practice it read nothing at all. Second, FindKey shared one forward-moving cursor across all lookups. The exporter writes /Subtype, /Rect, /Page, /Contents, /T, /Subj, but the importer searched in the order /Subtype, /Contents, /T, /Subj, /Page, /Rect; once the cursor had passed /Contents, the search for /Page and /Rect ran past the current entry and either found nothing or matched the keys of the next annotation. The library could not read its own output. Third, numbers went through PLStrToFloat, which follows the system decimal separator. ISO 32000-1 §12.7.7 defines FDF as PDF object syntax, and dictionary keys in PDF are unordered (§7.3.7), so any FDF parser that assumes a key order is wrong by construction, whichever tool produced the file

The repaired importer bounds each entry first. FindDictEnd walks from the opening << to its matching >>, tracking nested dictionaries and skipping literal string bodies with their backslash escapes, so a >> inside a comment such as (see section >> 4) cannot end the entry early. Every key lookup then starts at the entry's own start and is limited to its end, which makes key order irrelevant and stops one annotation from borrowing another's /Page. The key match also accepts a delimiter directly after the name, because /Contents(Hi) is as valid as /Contents (Hi), while the word-boundary rule keeps /Subj from matching the start of /Subtype and /T from matching /Type. ReadNumber now takes its position as a var parameter, skips whitespace and [, and parses with PLTryStrToFloatInvariant, which fails softly on a malformed token instead of raising. If any of the four rectangle numbers fails, all four fall back to zero rather than producing a half-read rectangle

PDFlibPas FindDictEnd now bounds each FDF annotation from its opening << to the matching >>, so every key lookup restarts at the entry start and stops at its end, and ReadNumber takes a var position, skips the bracket and parses with PLTryStrToFloatInvariant
The shared cursor could not read the library's own export: once it passed /Contents, the /Page and /Rect searches ran into the next annotation's keys, so key order is no longer allowed to matter

Why did FDF round-trips shift every annotation by its own height?

The old exporter wrote a rectangle in the wrong coordinate model. An annotation's /Rect is [llx lly urx ury] in default user space (ISO 32000-1 §12.5.2, with rectangles defined in §7.9.5), and FDF carries the same array. ExportAnnotationsToFDFString, however, called GetAnnotRectEx, which reports Left, Top, Width and Height in the library's drawing coordinates, the space that SetOrigin controls, and serialised them as [L T L+W T+H]. The importer, once it worked, wrote those four values back verbatim as a PDF rectangle, so the top edge landed where the lower-left corner belonged and every round trip moved the annotation up by its own height. The exporter now copies the annotation's own /Rect numbers, three decimals, dot separator, no exponent, and falls back to the computed rectangle only when the stored array is missing or not four numbers

PDFlibPas used to serialise FDF /Rect as left, top, width, height in drawing coordinates, so importing those four numbers back as llx lly urx ury landed the top edge where the lower-left corner belonged and moved every annotation up by its own height on each round trip
The exporter now copies the annotation's own /Rect numbers — three decimals, dot separator, no exponent — and the regression test compares a second export byte for byte with the first

The regression test that pins this down is worth copying, because it asserts on the document and on a second export, not on the importer's return value. Note the expected count of 2: AddNoteAnnotation creates a Text annotation plus its Popup, and both travel. The test also runs the export and import under a comma decimal separator, which is where the other half of this story lives

var
  Source, Target: TPDFlib;
  FDF: AnsiString;
  OldSep: Char;
begin
  Source := TPDFlib.Create;
  Target := TPDFlib.Create;
  try
    Source.NewPages(1);                     // now two pages
    Source.SelectPage(2);
    Source.AddNoteAnnotation(50.5, 60.25, 0, 80, 80, 120, 60,
      'Reviewer', 'Check this', 0.25, 0.5, 0.75, 0);
    Target.NewPages(1);

    OldSep := FormatSettings.DecimalSeparator;
    FormatSettings.DecimalSeparator := ',';   // simulate a German or French desktop
    try
      FDF := Source.ExportAnnotationsToFDFString;   // still writes /Rect [50.5 ...
      Target.ImportAnnotationsFromFDFString(FDF);
    finally
      FormatSettings.DecimalSeparator := OldSep;
    end;

    Target.SelectPage(2);
    Assert(Target.AnnotationCount = 2);           // the note and its popup
    Assert(Target.GetAnnotType(1) = 'Text');
    Assert(Target.ExportAnnotationsToFDFString = Source.ExportAnnotationsToFDFString);
  finally
    Target.Free;
    Source.Free;
  end;
end;

Be clear about what the FDF path carries. The importer rebuilds each entry as a dictionary with /Type, /Subtype, /Rect, /Contents, /T and /Subj; colour, flags, border style, popup links and appearance streams are not part of this route, and the exporter skips Widget annotations because form fields belong to the form-data methods. The broader map of which data travels through which method is in the overview of FDF, XFDF and XFA form data interchange, and if you need to inspect what actually arrived, the per-index readers such as GetAnnotType, GetAnnotTitle and GetAnnotContentsEx are covered in outline, annotation and action introspection

How do you read comma-decimal FDF and XFDF files from older exports?

For FDF the answer is unambiguous: a comma is not a delimiter in PDF syntax, so a number token that contains exactly one comma and no dot can only be a decimal written on a comma-locale machine. Earlier versions did write such files, for example /Rect [10,500 20,250 40,750 60,125], and the new ReadNumber turns that single comma into a dot before parsing. A token with two commas, or a comma and a dot, is rejected rather than guessed at. The reader does not consume exponent notation either, which matches ISO 32000-1 §7.3.3: PDF numbers never use it

XFDF is harder, because in XML attributes the comma is the separator. Standard XFDF (ISO 19444-1) writes rect="50.5,80.25,70.75,100.125" and dashes="4,2", while v3.539.28 and earlier, on a comma-locale system, wrote rect="50,500 80,250 70,750 100,125" and opacity="0,600", and also failed with EConvertError when reading a standard opacity="0.6". Since v3.539.29 both directions are invariant, and the legacy shape is recognised by XFDFNormalizeLegacyDecimals only when the attribute splits on whitespace into exactly the expected number of tokens (four for rect, one for opacity and width) and every token has the form digits-comma-digits. A standard rect never matches: it is either one token with three commas or tokens that end in a comma. dashes is deliberately left alone, because 4,2 could be two dash lengths or a legacy 4.2, and no rule can tell them apart

const
  // Keys out of exporter order, plus comma decimals from an older comma-locale export
  LegacyFDF: AnsiString = '%FDF-1.2'#10'1 0 obj'#10'<< /FDF << /Annots ['#10 +
    '<< /Rect [10,500 20,250 40,750 60,125] /Page 0 /Contents (First) ' +
    '/Subtype /Text /T (Alpha) /Type /Annot >>'#10 +
    '] >> >>'#10'endobj'#10'trailer'#10'<< /Root 1 0 R >>'#10'%%EOF'#10;
var
  Lib: TPDFlib;
begin
  Lib := TPDFlib.Create;               // a fresh document has one page
  try
    Lib.ImportAnnotationsFromFDFString(LegacyFDF);
    Assert(Lib.AnnotationCount = 1);
    Assert(Lib.GetAnnotTitle(1) = 'Alpha');
    // Re-exported as XFDF with dot decimals: rect="10.500 20.250 40.750 60.125"
    Writeln(Lib.ExportAnnotationsToXFDFString);
  finally
    Lib.Free;
  end;
end;

What should an annotation import test actually assert?

A useful import test asserts on the state of the target document, never only on what the importer says about itself. Nothing in the test suite checked AnnotationCount after an FDF import, and the return value, the only number anyone looked at, was the one number the bug left intact. Three assertions would have caught every defect described here: the annotation count on the expected page, one field read back through GetAnnotType or GetAnnotContentsEx, and a second export compared byte for byte with the first. The same discipline applies to any API that rewrites document structure in bulk, including the field consolidation described in merging duplicate form fields: check the resulting tree, not a returned total. The FDF and XFDF annotation methods, with their file and string variants, ship in the losLab PDF Library for Delphi and C++Builder, and v3.539.30 or later is the version to run if comments must survive the trip, v3.539.40 or later if the returned count must match what was added