PDF Library for Delphi (PDFlibPas) extracts the certificates inside a PDF signature with a pure DER walk over the CMS SignedData stored in /Contents, with no CryptoAPI involved. Since v3.539.10 every nested read is bounded by its parent element, the zero padding after the CMS is cut at the length the CMS declares, and object identifiers encode their combined first subidentifier in base-128. The bounds rule and the OID fix both replaced code that produced wrong answers without raising an error, and the padding rule keeps the stricter reader from rejecting real signatures
The read side matters more than it looks. Long-term validation tooling has to pull the signer certificate and its issuers out of an existing signature before it can fetch revocation data, an audit report has to say who signed, and a Lazarus build on Linux has no Windows message functions to lean on. A parser in that position rarely crashes on bad input. The failure mode that hurts is a certificate count that includes bytes belonging to a neighbour, a signer match made against the wrong field, or an OID that quietly turns into a different OID. A signature pipeline built on top of that reports confident nonsense
Reading the signer certificates out of a signed PDF
Five TPDFlib methods cover the read side, and all of them take InputFile, Password, FieldName: each call opens the file read-only, answers, and closes it again. GetSignatureEmbeddedCertificateCount and GetSignatureEmbeddedCertificateDER enumerate the certificates set in encoding order, GetSignatureSignerCertificateDER returns the certificate that produced a given SignerInfo, and GetSignatureCertificateChainLength / GetSignatureCertificateChainDER walk from that signer toward the furthest issuer the signature itself carries. Indexes are zero-based. Keep the results in AnsiString, which is why the library returns them that way: a DER blob routed through string or a TStrings goes through a character set conversion and comes back corrupted
uses
SysUtils, Classes, PDFlibrary;
procedure SaveDer(const FileName: string; const Der: AnsiString);
var
Fs: TFileStream;
begin
Fs := TFileStream.Create(FileName, fmCreate);
try
if Der <> '' then
Fs.WriteBuffer(Der[1], Length(Der));
finally
Fs.Free;
end;
end;
const
Src = 'contract-signed.pdf';
Field = 'Signature1';
var
Pdf: TPDFlib;
ChainLen, I: Integer;
SignerDer, LastDer: AnsiString;
begin
Pdf := TPDFlib.Create;
try
WriteLn('Certificates in the CMS: ',
Pdf.GetSignatureEmbeddedCertificateCount(Src, '', Field));
SignerDer := Pdf.GetSignatureSignerCertificateDER(Src, '', Field, 0);
if SignerDer = '' then
raise Exception.Create('signer certificate missing or not matched');
SaveDer('signer.cer', SignerDer);
ChainLen := Pdf.GetSignatureCertificateChainLength(Src, '', Field, 0);
for I := 0 to ChainLen - 1 do
SaveDer(Format('chain-%d.cer', [I]),
Pdf.GetSignatureCertificateChainDER(Src, '', Field, 0, I));
if ChainLen > 0 then
begin
LastDer := Pdf.GetSignatureCertificateChainDER(Src, '', Field, 0, ChainLen - 1);
WriteLn('Next issuer, if any: ', Pdf.GetCertificateIssuerURLs(LastDer));
end;
finally
Pdf.Free;
end;
end;
Two things in that output need care. A count of 0 is not a diagnosis: a missing field, a wrong password, a blob that is not DER, and a SignedData that simply omits the optional certificates set all come back as 0 or an empty string, so log the field name next to the number. And a chain that ends before a self-issued certificate is not an error either. The chain builder only uses certificates embedded in the signature, so the remaining issuers have to be fetched through the addresses GetCertificateIssuerURLs reports
How much of /Contents is actually the CMS?
Only the prefix that the outer SEQUENCE declares belongs to the CMS, and PLTrimCMSPadding cuts everything after it. A signer reserves the /Contents hex string before the CMS exists, because the /ByteRange described in ISO 32000-1 §12.8.1 has to be fixed first, so the slot is sized generously and the unused tail is zeros. PLTrimCMSPadding reads the first TLV, requires tag $30, and returns the bytes up to the end of that element; anything that does not start with a well-formed SEQUENCE comes back empty. That top level is the one place where trailing bytes are legal, and the distinction matters for the next section: a strict "the element must consume the whole buffer" rule would reject every real-world signature, while a lax rule applied at every depth lets nested fields read bytes they do not own
Why does a DER reader need the parent's end offset?
A nested element is only valid if it ends inside its parent, and checking against the end of the buffer does not prove that. The low-level DERReadTLV in PDFlibASN1 bounds each element against the whole string, which is the right check for the outermost object and the wrong one for everything below it. Picture a SignerInfo whose issuerAndSerialNumber declares 40 bytes while the issuer Name inside it claims 60. Every byte is still in the buffer, so a buffer-bounded reader accepts the Name, reads the serial number out of the digest algorithm that follows, and then compares that pair against the embedded certificates. Before v3.539.10 the CMS walker read exactly that way. The fix is a small wrapper that carries the parent's end position into every read
function ReadTLVWithin(const Data: AnsiString; ParentEnd: Integer;
var Offset: Integer; out Tag: Byte; out Start, Len: Integer): Boolean;
begin
Result := False;
// nothing left inside the parent: refuse to start a read
if (Offset < 1) or (Offset >= ParentEnd) then
Exit;
if not DERReadTLV(Data, Offset, Tag, Start, Len) then
Exit;
// Offset now sits one past the element; it must not pass the parent
Result := Offset <= ParentEnd;
end;
// each level records its own end and hands it down:
// OuterEnd := end of ContentInfo (RFC 5652 section 3)
// ExplicitEnd := end of content [0] EXPLICIT
// ContentEnd := end of SignedData (RFC 5652 section 5.1)
// SignerEnd / InnerEnd for SignerInfo and issuerAndSerialNumber
The unit, PDFlibCMSRead, now threads those ends through ContentInfo, the [0] EXPLICIT wrapper, the SignedData fields up to signerInfos, the SignerIdentifier in both its issuerAndSerialNumber and [0] subjectKeyIdentifier forms (RFC 5652 §5.3), and the tbsCertificate fields read from each embedded certificate when matching the signer. Inside the certificates set and the signerInfos set, an element that runs past the set's end stops the loop: PLExtractCMSCertificates returns the certificates it had already accepted and never glues the following crls or signerInfos bytes onto the last one. The issuer-and-serial match also requires both halves, since a serial number is only unique within one issuer
Why did 2.999.3 come out as 1.15.3?
The first two arcs of an OID combine into one subidentifier, not one byte, and that subidentifier is base-128 encoded like every other arc. X.690 §8.19.4 defines it as 40 * arc1 + arc2; the earlier DER_OID wrote that value with Byte(...), which is correct only up to 127, the value of 2.47. For 2.999 the sum is 1079, the byte cast keeps 55, and 55 decodes as 1.15, so the identifier silently names a different branch of the tree. Values from 128 to 255 fail differently, emitting one byte with the continuation bit set that swallows the next arc. Most PKI identifiers (1.2.840..., 2.5.29..., 0.4.0...) never reach the boundary, which is why it survived; the joint-iso-itu-t arcs from 2.48 upward do. DER_OID serves both the encoder for signed attributes and the matcher in DERFindExtensionByOID and the SignedData content-type check, so a wrong encoding broke writing and lookup alike
uses
SysUtils, PDFlibASN1;
function Hex(const S: AnsiString): string;
var
I: Integer;
begin
Result := '';
for I := 1 to Length(S) do
Result := Result + IntToHex(Byte(S[I]), 2) + ' ';
Result := Trim(Result);
end;
begin
WriteLn(Hex(DER_OID('2.999.3'))); // 06 03 88 37 03
WriteLn(Hex(DER_OID('2.47.1'))); // 06 02 7F 01
WriteLn(Hex(DER_OID('2.48.1'))); // 06 03 81 00 01
WriteLn(Hex(DER_OID('2.5.29.14'))); // 06 03 55 1D 0E
WriteLn(Hex(DER_OID('1.2.840.113549.1.7.2'))); // 06 09 2A 86 48 86 F7 0D 01 07 02
end.
The combined value is held in a UInt64 on purpose. DER_OID parses arcs into Int64, so a legal second arc can be as large as Int64.MaxValue, and adding 80 for arc1 = 2 overflows a signed 64-bit integer. UInt64 carries Int64.MaxValue + 80 without wrapping, and the ten-byte scratch buffer holds the ten 7-bit groups a 64-bit value needs. Test vectors worth keeping are the ones on either side of the boundary: 2.47 must stay one byte, and 2.48 must become two
What does the read-side CMS walker guarantee?
PDFlibCMSRead guarantees structure and nothing else: it returns bytes that sit where RFC 5652 says they should and verifies no signature, digest, or validity period. The walker accepts DER only, so DERReadTLV rejects indefinite lengths and multi-byte tag numbers, and a BER-encoded CMS from a nonconforming signer reports zero certificates rather than a partial guess. Attribute certificates and the other CertificateChoices alternatives are skipped because nothing downstream can use them. Cryptographic verification stays with the code that owns it, which starts with the byte coverage checks described in PAdES signing and ByteRange validation in Delphi and continues with classifying what changed after a PDF was signed
The broader lesson carries to any binary format: "inside the buffer" is a memory-safety property, "inside the parent" is a correctness property, and a parser needs both. The same thinking about hostile lengths runs through hardening a Pascal PDF parser against malicious files. The certificate extraction, chain building, and long-term validation APIs discussed here ship with losLab PDF Library for Delphi, for Delphi, C++Builder and Lazarus