Technical Article

PDF Number Grammar vs JSON: NaN, Infinity and null in Delphi

PDF Library for Delphi (PDFlibPas) emits valid JSON for every PDF number since v3.539.31. GetObjectJSON rewrites tokens that ISO 32000-1 accepts but RFC 8259 rejects, such as -.25, +1.5 and 007.5, into -0.25, 1.5 and 7.5 digit for digit; GetDocumentJSON and the analysis reports write null for NaN and Infinity; and PLDoubleToStr writes 0 for NaN instead of raising EInvalidOp halfway through an export. Before the fix, the library could produce JSON that its own reader refused to load back

Why does a valid PDF number break JSON?

Because the two grammars disagree on four small details, and a PDF parser that respects the source text will carry those details straight into the output. ISO 32000-1 §7.3.3 lets a number start with a plus sign, omit the integer part (.5), end on a bare point (4.) and carry leading zeros (007.5). RFC 8259 §6 allows none of that: an optional minus, an integer part that is either 0 or starts with 1 to 9, and at least one digit after any decimal point. Producers are free to write the PDF forms, and plenty of generators and hand-edited files do

The leak came from a deliberate precision feature. Since v3.539.19, TPDFNumeric.Output returns the exact text the tokenizer parsed for real numbers, which is what keeps a calibrated colour value exact on save, as described in preserving parsed PDF decimal precision. The tokenizer already patches .5 into 0.5 and 4. into 4.0 on the way in, and integers are reformatted from their value, so +3 comes back as 3. What survives verbatim is the rest: a signed leading point (-.25), an explicit plus on a real (+1.5) and leading zeros (007.5). The old object writer appended Output right after "value":, and TJSONParser.ParseNumber in the library's own reader stops on every one of those with "Invalid JSON number", so the export succeeded and the re-import failed with PDFLIB_ERROR_OBJECT_JSON_INVALID (105)

PDFlibPas GetObjectJSON rewrites the PDF number tokens RFC 8259 rejects digit for digit: -.25 becomes -0.25, +1.5 loses the plus, 007.5 drops its leading zeros, and fractional digits like 1.250000 survive, because formatting from the stored Double would add binary noise
The old writer appended the exact parsed text, the library's own reader stopped with Invalid JSON number, and error 105 broke a round trip the export side called a success
uses
  System.SysUtils, PDFlibrary;

var
  Lib: TPDFlib;
  JSON: AnsiString;
begin
  Lib := TPDFlib.Create;
  try
    if Lib.LoadFromFile('legacy-drawing.pdf', '') = 0 then
      raise Exception.Create('load failed');

    // Object 12 is an array written as [-.25 +1.5 007.5]
    JSON := Lib.GetObjectJSON(12, 0);
    // v3.539.31 and later: the values arrive as -0.25, 1.5 and 7.5

    // SetObjectJSON accepts no options, so pass 0
    if Lib.SetObjectJSON(12, JSON, 0) = 0 then
      raise Exception.CreateFmt('round trip rejected, error %d',
        [Lib.LastErrorCode]);

    Lib.SaveToFile('legacy-drawing-roundtrip.pdf');
  finally
    Lib.Free;
  end;
end;

How does PDFNumberTextToJSON keep every digit?

PDFNumberTextToJSON re-spells the token instead of recomputing it from a Double. The function in PDFlibObjectJSON reads an optional sign, collects digits before and after a single decimal point, and then applies only the edits JSON demands: it drops a plus, strips leading zeros while keeping one, supplies 0 when the integer part is empty, drops a bare trailing point and puts the minus back. A token containing any other character, or no digits at all, falls back to PLJSONNumber(Value, 10), which writes null when the value is not finite

PDFlibPas PDFNumberTextToJSON reads the sign, collects digits around a single decimal point and applies only the edits JSON demands, while any other character or an empty digit run falls back to PLJSONNumber, which writes null for NaN and Infinity instead of a number
Re-spelling beats recomputing: the tokenizer already patched .5 and 4. on the way in, so the writer keeps every surviving digit and the round trip recreates exactly the same value
  • -.25 becomes -0.25, and +.5 becomes 0.5
  • +1.5 becomes 1.5
  • 007.5 becomes 7.5, while 0.75 stays as it is
  • 4. becomes 4 if such a token ever reaches the writer
  • 2.22221 and 1.250000 keep every fractional digit, trailing zeros included

Formatting from the stored Double would have been shorter and wrong, for the same reason the precision fix exists: the default output precision is four decimals, and even a full-precision conversion can add binary noise to a decimal literal. Keeping the digits means SetObjectJSON and ImportObjectJSON, which hand each JSON number text to the PDF tokenizer, recreate exactly the same value. The guarantee covers the value, not the bytes: after a re-import, -.25 is stored and saved as -0.25. Both spellings are equal under §7.3.3, but a byte-level diff will flag the change, so do not treat an export and import cycle as a no-op on a document whose bytes are covered by a signature

What happens to a number JSON cannot represent?

GetDocumentJSON now writes null for any number that is NaN or infinite, because RFC 8259 §6 has no syntax for either. Infinity is easier to produce than it sounds: the PDF tokenizer accumulates digits by repeated multiplication into a Double, which tops out near 1.8 × 10308, so an integer literal a little over 300 digits long silently becomes +Inf. Honest files never contain such a literal; fuzzed and hostile ones do, which is why they belong in the same test corpus as the cases in hardening a Pascal PDF parser against malicious files. The old document writer formatted non-integers with Str(D:0:6), and for +Inf that writes the text +Inf, which no JSON consumer will parse

The null is deliberately lossy. Consumers of GetDocumentJSON output must accept null anywhere a number can appear, and should read it as "a value was present but cannot be represented", not as a missing key. The original literal is not recoverable from the document JSON, so a pipeline that cares should log the object and treat the file as suspect rather than substitute a default

Why could a single NaN abort an SVG or JSON export?

Because PLDoubleToStr, the invariant number formatter behind content streams, SVG, XML, CSV and most JSON in the library, scaled its input and called Round, and Round(NaN) raises EInvalidOp on targets such as Win32, where Delphi leaves the x87 invalid-operation exception unmasked. The exception fired after the writer had already emitted part of its output, so one degenerate measurement, a 0/0 in a metric or a NaN passed in by a caller, left a truncated file behind. PLDoubleToStr now returns 0 for NaN, and its integer branch clamps to ±9.2e18 like the fractional branch, so Infinity also comes out as a finite literal

Zero is the right answer for a content stream, where a number slot must hold a number, and the wrong answer for a report, where 0 is a plausible measurement. JSON writers that must keep the difference use PLJSONNumber(Value, Decimals) from PDFlibExtra, which writes null for NaN or Infinity and invariant digits otherwise. PLJSONNumber now backs GetSimilarImageDeduplicationReportJSON, GetAnnotationHitsJSON and the barcode, deskew, structured text and PDF/VCR reports; the deskew report previously wrote 0 for a non-finite angle and now writes null

PDFlibPas stops NaN and Infinity three ways: AddPageMatrix, ScalePage and RedactRegion reject non-finite arguments up front, PLDoubleToStr writes 0 for content-stream slots, and PLJSONNumber writes null in reports, where zero would read as a plausible measurement, after Round(NaN) used to raise EInvalidOp mid-export
Zero is the right answer for a content stream and the wrong answer for a report, so the report writers hand every Double to PLJSONNumber and let null say the value was present but unrepresentable
uses
  SysUtils, PDFlibTypes, PDFlibExtra;

function SkewReportJSON(Page: Integer; Angle, Confidence: Double): string;
var
  B: PLStringBuilder;
begin
  B := PLStringBuilder.Create(128);
  try
    // Format every Double to text first; PLJSONNumber writes null
    // for NaN or Infinity and always uses a decimal point
    B.Append('{"page":').Append(Page)
     .Append(',"angle":').Append(string(PLJSONNumber(Angle, 4)))
     .Append(',"confidence":').Append(string(PLJSONNumber(Confidence, 4)))
     .Append('}');
    // Never B.Append(Angle): the Double overload follows the user locale
    Result := B.ToString;
  finally
    B.Free;
  end;
end;

Where does the user locale still sneak into JSON?

Through any formatter that consults regional settings, and a full audit of machine-readable output found exactly one left: maxAcceptedMeanError in GetSimilarImageDeduplicationReportJSON, written with PLFloatToStr, a thin wrapper over FloatToStr. On a desktop whose decimal separator is a comma, the report contained "maxAcceptedMeanError":1,5, which a JSON parser reads as the value 1 followed by a stray token. The field reports the worst accepted pixel error from perceptual image deduplication, and it now goes through PLJSONNumber(Stats.MaxAcceptedMeanError, 6). A remaining trap is PLStringBuilder: on Delphi it is a plain alias of System.SysUtils.TStringBuilder, whose Append(Double) overload formats through the user locale, while FPC builds use a library class instead, so a test on Free Pascal or an en-US machine will never catch it

uses
  System.SysUtils, System.JSON, PDFlibrary;

var
  Lib: TPDFlib;
  Report: WideString;
  Parsed: TJSONValue;
begin
  // Reproduce a German or French desktop inside the test run
  FormatSettings.DecimalSeparator := ',';
  Lib := TPDFlib.Create;
  try
    // Use a fixture that really contains near-duplicate images,
    // otherwise the mean error is 0 and the bug stays hidden
    Lib.LoadFromFile('scanned-batch.pdf', '');
    // Dry run with thresholds 2, 2, 4: the document is not modified
    Lib.GetSimilarImageDeduplicationReportJSON(2, 2, 4, Report);
    Parsed := TJSONObject.ParseJSONValue(Report);
    if Parsed = nil then
      raise Exception.Create('report is not valid JSON on a comma locale');
    Parsed.Free;
  finally
    Lib.Free;
  end;
end;

A regression suite for JSON output needs three fixtures to stay honest: a page carrying -.25, +1.5 and 007.5, an object holding a 400-digit integer, and any report run under a comma locale, each validated with a strict parser rather than eyeballed. Object JSON, document JSON and the analysis reports in PDF Library for Delphi share the same number rules across Delphi, C++Builder and Free Pascal; the complete feature list is on the PDF Library for Delphi product page