Technical Article

HotPDF OCR Text Layer Offset: Mapping Through the CropBox

OCR text layers, barcode bounds and face redaction boxes drift on cropped PDF pages when bitmap pixels are mapped back through the MediaBox instead of the box the renderer actually rasterized: the CropBox clipped to the MediaBox (ISO 32000-1 §14.11.2). HotPDF fixed this for ApplyLoadedOCRTextLayer in v2.770.153, and for DecodeLoadedPageBarcodes and DetectLoadedRedactionFindings in v2.770.154

The bug report that usually arrives looks like this. A scanned contract archive goes through OCR, the output is searchable, and the search hit for a clause number is highlighted half an inch below and to the left of the printed number. Most files in the batch are fine. The broken ones all came from one scanning station that writes a /CropBox to trim the platen margin. That one detail separates the picture the OCR engine saw from the frame the text layer was placed in, and the same mismatch moves barcode bounds and, more seriously, face redaction boxes

Why does the OCR text layer drift away from the scanned words?

The text layer drifts because two halves of the pipeline disagreed about which rectangle the bitmap covers. In v2.766.64, HotPDF changed rendering, SVG export, the viewer and printing to honor the CropBox: a page is displayed through its CropBox clipped to its MediaBox, which is what ISO 32000-1 §14.11.2 prescribes, and GetLoadedPageVisibleBox was added to return that visible box. The recognition features kept building their device-to-page transform from GetLoadedPageBox(PageIndex, pbMediaBox, ...). The raster now covered the visible box, the transform still assumed the MediaBox, and every recognized position came back offset by the gap between the two

The affected window is therefore precise. ApplyLoadedOCRTextLayer misplaced text from v2.766.64 through v2.770.152. Whole-page DecodeLoadedPageBarcodes and the face detection inside DetectLoadedRedactionFindings stayed wrong one build longer, through v2.770.153. Before v2.766.64 the renderer drew the whole MediaBox, so mapping and raster agreed, at the cost of recognizing content that viewers never show. The fixes changed three things together for each feature: the transform, the pixel budget estimate, and the page box handed to a custom engine in the request record

Several cases were never affected:

  • Pages without a /CropBox, or whose CropBox equals the MediaBox, map identically before and after the fix
  • DecodeLoadedPageBarcodes with HasRegion set renders exactly the region you pass and maps through that same region, so explicit-region decoding was correct throughout; the check that the region lies inside the page still uses the MediaBox
  • Pattern-based redaction findings (emails, card numbers and so on) come from text extraction in user space, not from a raster, so only the face-detection findings moved

Three coordinate frames, and which HotPDF APIs use each

HotPDF code that touches recognition deals with three frames, and most mapping bugs come from mixing two of them

  • Bitmap pixels: top-left origin, Y grows downward, units are pixels at the request DPI. THPDFOCRWord.Left, Top, Right and Bottom are in this frame, as are the optional baseline points, the results a custom IHPDFBarcodeDecoder returns, and the boxes from a custom IHPDFFaceDetector
  • PDF user space of a loaded page: bottom-left origin, Y grows upward, units are points, with Bottom < Top. GetLoadedPageBox and GetLoadedPageVisibleBox return Left, Bottom, Right, Top in this frame, and so do the PageLeft, PageBottom, PageRight and PageTop fields of THPDFOCRRequest, the bounds in THPDFDecodedBarcode and the rectangles in THPDFRedactionFinding
  • HotPDF page-drawing coordinates: the API you use to build new pages (text output, shapes, barcodes, links, form fields) works with a top-left origin and Y growing downward. That frame belongs to document generation and has nothing to do with the loaded-document APIs above, so never feed a loaded-page user-space rectangle into it unchanged

The OCR word record is deliberately pixel-based: an engine reports what it saw in the image, and ApplyLoadedOCRTextLayer owns the conversion. That division only works when the conversion uses the right box, which is what v2.770.153 restored

HotPDF recognition coordinate frames: bitmap pixels with a top-left origin used by THPDFOCRWord boxes and custom decoders, PDF user space with a bottom-left origin returned by GetLoadedPageBox and GetLoadedPageVisibleBox, and the top-left page drawing API, which must never receive a loaded-page rectangle unchanged
engines report pixels because that is what they saw, HotPDF maps them, and mixing the two frames is how layers and redaction boxes drift

The device-to-page transform behind OCR, barcodes and faces

HotPDF maps bitmap pixels to the page with a single affine matrix built from five inputs: the rotation, the scale DPI / 72, the bitmap height, and the Left, Bottom, Right and Top of the rendered box. OCR, barcode decoding and face detection all share one routine for it, which is why one wrong box input broke all three in the same way. For an unrotated page the page-to-device matrix [A B C D E F] is:

  • A = Scale and D = -Scale, where Scale = DPI / 72; the negative D flips user space (Y up) into bitmap space (Y down)
  • B = C = 0, because an unrotated page has no shear or swap between axes
  • E = -Left * Scale, which moves the box's left edge to pixel column 0
  • F = BitmapHeight + Bottom * Scale, which maps the box's bottom edge to y = BitmapHeight, the lower edge of the bitmap, so the top edge lands on row 0

Pixels go back to the page through the inverse of that matrix. Request.PageRotation carries the page's /Rotate normalized to 0, 90, 180 or 270 (any value that is not a multiple of 90 is treated as 0), and the renderer turns the page clockwise as ISO 32000-1 §7.7.3.3 requires. Under rotation the axes swap and a different pair of box edges is pinned to the bitmap origin. Written out as inverse formulas, with S = DPI / 72, x and y in pixels and H the bitmap height:

/RotatePage XPage YBox edges the mapping depends on
0Left + x / SBottom + (H - y) / SLeft, Bottom
90Left + y / SBottom + x / SLeft, Bottom
180Right - x / SBottom + y / SRight, Bottom
270Right - y / STop - x / SRight, Top

The last column explains why the bug looked random in production. A CropBox that trims only the top of the page leaves Left and Bottom untouched, so upright pages came out perfectly and only pages carrying /Rotate 270 drifted. Rotation also swaps the bitmap dimensions: at 90 and 270 the bitmap is (Top - Bottom) * S pixels wide and (Right - Left) * S pixels tall

HotPDF inverse mapping table for page rotation: at /Rotate 0 and 90 the transform pins the Left and Bottom edges of the rendered box, at 180 it pins Right and Bottom, at 270 Right and Top, which is why a cropped page drifts in a different direction for every orientation in a mixed document
the same half-inch crop looks like three different bugs once pages carry different /Rotate values, because each orientation pins a different pair of box edges

What goes wrong with MediaBox [0 0 612 792] and CropBox [36 36 576 756]?

With a half-inch crop on every side, an unrotated page's text layer lands exactly 36 points left of and 36 points below the scanned words when the MediaBox is used. Take a US Letter page whose CropBox trims 36 points (0.5 inch) from each edge. The visible box is 540 by 720 points, so at the default OCR resolution of 300 DPI the scale is 300 / 72 ≈ 4.1667 and the bitmap is 2250 by 3000 pixels

Suppose the engine reports a word with pixel box Left 450, Top 600, Right 900, Bottom 660 and no baseline. HotPDF then places the baseline 20 percent of the word height above the bottom edge, at pixel row 648, and maps the start point (450, 648):

  • Through the visible box: x = 36 + 450 / 4.1667 = 144.0 and y = 36 + (3000 - 648) / 4.1667 = 600.48, which is where the word is printed
  • Through the MediaBox: x = 0 + 108.0 = 108.0 and y = 0 + 564.48 = 564.48, a uniform shift of (-36, -36) points
HotPDF CropBox drift anatomy on a US Letter page with MediaBox 0 0 612 792 and CropBox 36 36 576 756: the renderer rasterizes the visible box at 300 DPI, so mapping word pixel 450 through GetLoadedPageVisibleBox gives 144.0 and 600.48 while the MediaBox transform lands at 108.0 and 564.48
the raster covers the CropBox, so any transform built from the MediaBox shifts every recognized word by exactly the crop margin

Rotate the same page and the direction of the error changes, because different edges are involved. At /Rotate 180 the X term uses Right, and 612 instead of 576 pushes the layer 36 points to the right while Bottom still pulls it 36 points down. At /Rotate 270 both Right and Top are too large, so the layer moves 36 points right and 36 points up. A document with mixed orientations can show the drift in three directions, a reliable fingerprint for this bug. Hand-written code that derives the scale from the box, such as Bitmap.Width / (Right - Left), also stretches every coordinate by 612 / 540, roughly 13 percent, on top of the offset

Which of your PDF documents are affected?

A PDF document is exposed when at least one page has a visible box that differs from its MediaBox, and HotPDF can tell you that in a few lines. Compare GetLoadedPageBox with pbMediaBox against GetLoadedPageVisibleBox for every page, and print GetLoadedPageRotation alongside so you can predict the drift direction from the table above. THPDFPageBoundary also offers pbCropBox, pbBleedBox, pbTrimBox and pbArtBox, but GetLoadedPageBox(pbCropBox) falls back to the MediaBox when no crop box exists and does not clip, so the visible box is the right thing to compare against

uses
  System.SysUtils, HPDFDoc;

procedure ReportCroppedPages(const FileName: string);
var
  Pdf: THotPDF;
  I: Integer;
  ML, MB, MR, MT, VL, VB, VR, VT, Tmp: Single;
begin
  Pdf := THotPDF.Create(nil);
  try
    if Pdf.LoadFromFile(FileName) < 1 then
      raise Exception.Create('Cannot load ' + FileName);
    for I := 0 to Pdf.LoadedPageCount - 1 do
    begin
      if not Pdf.GetLoadedPageBox(I, pbMediaBox, ML, MB, MR, MT) then
        Continue;
      // the stored array may list its corners in any order
      if MR < ML then begin Tmp := ML; ML := MR; MR := Tmp; end;
      if MT < MB then begin Tmp := MB; MB := MT; MT := Tmp; end;
      // already normalized and clipped to the MediaBox
      if not Pdf.GetLoadedPageVisibleBox(I, VL, VB, VR, VT) then
        Continue;
      if (Abs(VL - ML) > 0.01) or (Abs(VB - MB) > 0.01) or
         (Abs(VR - MR) > 0.01) or (Abs(VT - MT) > 0.01) then
        Writeln(Format('Page %d  MediaBox [%g %g %g %g]  visible [%g %g %g %g]  /Rotate %d',
          [I + 1, ML, MB, MR, MT, VL, VB, VR, VT,
           Pdf.GetLoadedPageRotation(I)]));
    end;
  finally
    Pdf.Free;
  end;
end;

Two details of GetLoadedPageVisibleBox matter for scripts like this. The function leaves its out parameters untouched when it fails, so presetting a default page size before the call is a safe pattern. And when a malformed CropBox does not intersect the MediaBox at all, the function returns the MediaBox rather than an empty rectangle. If the report lists pages and your deployed build is older than v2.770.153 for OCR, or v2.770.154 for barcodes and faces, re-run recognition on those pages after upgrading. An OCR layer committed by an affected build stays in the saved file, and the default SkipPagesWithText option will skip those pages on a second pass unless you turn it off or remove the old layer first

How should a custom IHPDFOCREngine map pixels back to PDF space?

A custom IHPDFOCREngine should return word boxes in bitmap pixels and let HotPDF do the mapping; convert to user space only for your own decisions, and then use the box from the request, never the MediaBox. Since v2.770.153 the request's PageLeft, PageBottom, PageRight and PageTop describe the rendered visible box, so they match Request.Bitmap exactly. The helper below is the inverse of the library's transform, including its use of the actual bitmap height for upright pages, so it agrees with HotPDF to the pixel

uses
  System.SysUtils, System.Math, Vcl.Graphics, HPDFDoc;

// Bitmap pixel (top-left origin, Y down) to PDF user space
// (bottom-left origin, Y up), through the box the bitmap was rendered from
procedure HotPixelToPage(Rotation, DPI, BitmapHeight: Integer;
  Left, Bottom, Right, Top: Single; X, Y: Double;
  out PageX, PageY: Double);
var
  S: Double;
begin
  S := DPI / 72.0;
  case Rotation of
    90:  begin PageX := Left + Y / S;  PageY := Bottom + X / S; end;
    180: begin PageX := Right - X / S; PageY := Bottom + Y / S; end;
    270: begin PageX := Right - Y / S; PageY := Top - X / S; end;
  else
    PageX := Left + X / S;
    PageY := Bottom + (BitmapHeight - Y) / S;
  end;
end;

A realistic reason to need user space inside an engine is a zone rule: invoices whose letterhead you never want searchable, or a stamp area that confuses the recognizer. The engine below, written with TInterfacedObject so reference counting handles its lifetime, filters words by where their centers fall on the page, then returns the survivors untouched in pixel coordinates. RunRecognizer stands in for your own recognizer call

type
  TZoneFilterOCREngine = class(TInterfacedObject, IHPDFOCREngine)
  private
    FSkipLeft, FSkipBottom, FSkipRight, FSkipTop: Single;  // user space
    function RunRecognizer(Bitmap: TBitmap; MaxWords: Integer;
      out Words: THPDFOCRWords): boolean;  // your recognizer, pixel boxes
  public
    constructor Create(SkipLeft, SkipBottom, SkipRight, SkipTop: Single);
    function GetName: AnsiString;
    function Recognize(const Request: THPDFOCRRequest;
      out Words: THPDFOCRWords; out Diagnostic: AnsiString): boolean;
  end;

function TZoneFilterOCREngine.Recognize(const Request: THPDFOCRRequest;
  out Words: THPDFOCRWords; out Diagnostic: AnsiString): boolean;
var
  Raw: THPDFOCRWords;
  I, Count: Integer;
  CX, CY: Double;
begin
  Diagnostic := '';
  SetLength(Words, 0);
  if not RunRecognizer(Request.Bitmap, Request.MaxWords, Raw) then
  begin
    Diagnostic := 'recognizer failed';
    Exit(False);
  end;
  SetLength(Words, Length(Raw));
  Count := 0;
  for I := 0 to High(Raw) do
  begin
    HotPixelToPage(Request.PageRotation, Request.DPI,
      Request.Bitmap.Height, Request.PageLeft, Request.PageBottom,
      Request.PageRight, Request.PageTop,
      (Raw[I].Left + Raw[I].Right) / 2, (Raw[I].Top + Raw[I].Bottom) / 2,
      CX, CY);
    if (CX >= FSkipLeft) and (CX <= FSkipRight) and
       (CY >= FSkipBottom) and (CY <= FSkipTop) then
      Continue;
    Words[Count] := Raw[I];  // still pixels: HotPDF maps them itself
    Inc(Count);
  end;
  SetLength(Words, Count);
  Result := True;
end;

Hand the engine to ApplyLoadedOCRTextLayer(PageIndices, Engine, Options, Info) as with any other engine. The library validates what comes back before it trusts it: a word is dropped and counted in Info.DroppedWordCount when its box leaves the bitmap, when Right <= Left or Bottom <= Top, or when Confidence is outside 0..1 or below MinimumConfidence. Returning more words than MaxWordsPerPage, or pushing the running total past MaxTotalWords, fails the whole call with a budget error, so honor Request.MaxWords in the engine. Do not convert word boxes to user space before returning them; HotPDF would treat the point values as pixels and the layer would collapse toward the bitmap origin

Mapping your own detector output

The same helper serves a home-grown pipeline built on RenderLoadedPageToBitmap, which renders the visible box and applies /Rotate just as the recognition features do. Read the box with GetLoadedPageVisibleBox, normalize the rotation the same way HotPDF does, and map two opposite corners of each pixel box. The Y axis flips and, at 90 and 270 degrees, the axes swap, so the mapped corners come out in no fixed order; take the minimum and maximum of the mapped points, which is also how HotPDF builds barcode bounds

const
  DPI = 200;
var
  Pdf: THotPDF;
  Bmp: TBitmap;
  VL, VB, VR, VT: Single;
  Rotation: Integer;
  PxL, PxT, PxR, PxB, X1, Y1, X2, Y2: Double;
begin
  Pdf := THotPDF.Create(nil);
  try
    Pdf.LoadFromFile('scanned-ids.pdf');
    if not Pdf.GetLoadedPageVisibleBox(0, VL, VB, VR, VT) then Exit;
    Rotation := Pdf.GetLoadedPageRotation(0) mod 360;
    if Rotation < 0 then Inc(Rotation, 360);
    if (Rotation <> 90) and (Rotation <> 180) and (Rotation <> 270) then
      Rotation := 0;
    Bmp := Pdf.RenderLoadedPageToBitmap(0, DPI);
    if Bmp = nil then Exit;
    try
      MyDetector(Bmp, PxL, PxT, PxR, PxB);  // your code, pixel box
      HotPixelToPage(Rotation, DPI, Bmp.Height, VL, VB, VR, VT,
        PxL, PxT, X1, Y1);
      HotPixelToPage(Rotation, DPI, Bmp.Height, VL, VB, VR, VT,
        PxR, PxB, X2, Y2);
      Writeln(Format('User-space box [%.2f %.2f %.2f %.2f]',
        [Min(X1, X2), Min(Y1, Y2), Max(X1, X2), Max(Y1, Y2)]));
    finally
      Bmp.Free;
    end;
  finally
    Pdf.Free;
  end;
end;

The rotation behavior is covered in more depth in flattening page rotation without breaking the page boxes, and the barcode decoding pipeline that consumes the same transform in decoding rotated QR codes from PDF pages. If your engine wraps an external recognizer, the Tesseract OCR adapter for searchable PDF shows the process isolation and cancellation side of the same interface

Quick reference: CropBox-safe coordinate mapping

  • The renderer rasterizes the visible box, the CropBox clipped to the MediaBox (ISO 32000-1 §14.11.2); every pixel-to-page mapping must use that box, read with GetLoadedPageVisibleBox
  • HotPDF v2.770.153 fixed ApplyLoadedOCRTextLayer; v2.770.154 fixed whole-page DecodeLoadedPageBarcodes and face findings from DetectLoadedRedactionFindings; builds from v2.766.64 up to those versions are affected
  • THPDFOCRWord boxes are bitmap pixels with a top-left origin; GetLoadedPageBox and GetLoadedPageVisibleBox return PDF user space with a bottom-left origin and Bottom < Top
  • Scale is DPI / 72; derive it from the DPI, never from a page box divided into the bitmap width
  • /Rotate decides which edges matter: Left and Bottom at 0 and 90, Right and Bottom at 180, Right and Top at 270
  • Return OCR words in pixels and let HotPDF map them; convert only for your own filtering logic
  • Re-run OCR on cropped pages processed by an affected build, and remember that SkipPagesWithText skips pages that already carry the old layer

The recognition features, page-box queries and loaded-document rendering used here all ship in the HotPDF component for Delphi and C++Builder; licensing, trial downloads and the full feature list are on the HotPDF Delphi PDF component page