OCR Components and Deployment
Estimated Reading Time: 2 MinutesThis article applies to the Windows version of Datalogics Adobe PDF Library (APDFL) 21. Feature availability, component filenames, file sizes, and deployment requirements may differ on other platforms or in other APDFL versions.
Description
Optical character recognition (OCR) recognizes text in scanned documents and images. It can add searchable text to PDFs whose pages would otherwise contain only images.
OCR makes scanned content easier to search, select, copy, and extract. It is useful for digitizing paper records, organizing document collections, and making scanned information available for subsequent processing.
Samples
The installed samples are:
- CPlusPlus\Sample_Source\OCR\OCRPage\OCRPage.cpp: recognizes text on a PDF page.
- CPlusPlus\Sample_Source\OCR\OCRImage\OCRImage.cpp: recognizes text in an image and produces PDF output.
Public samples (github):
OCRPage
OCRImage
Files
Paths are relative to the SDK root. Sizes are uncompressed file lengths, rounded to two decimal places. Sizes and savings may vary by SDK build.
| File or directory | Size (MB) |
| CPlusPlus\Binaries\DL210OCREngine.ppi | 5.99 |
| CPlusPlus\Binaries\dltesseract5.dll | 11.80 |
| CPlusPlus\Binaries\tessdata4\ (entire directory) | 860.04 |
| Total | 877.83 |
These components provide the OCR plug-in, recognition engine, and supporting recognition data. The tessdata4 directory total includes all supplied files and subdirectories.
Omitting the Feature When Unused
If OCR is unused, omit DL210OCREngine.ppi, dltesseract5.dll, and the entire tessdata4 directory. The combined saving is 877.83 MB.
PDF processing does not require OCR simply because the input contains scanned pages. Without the OCR components, tested PDF opening, rendering, saving, reopening, text extraction, and library shutdown succeeded. Scanned pages remained image-only. Text previously added by OCR remained extractable without the OCR engine, with matching rendering and word counts.
The positive OCR test added 18 recognized words to the supplied OCRPage input. Tests also covered omission of the plug-in, engine DLL, and language-data directory individually and together.
Application considerations
Remove unconditional OCR initialization and recognition calls when OCR is omitted. For optional installation, check gOCREngineHFT after InitOCREngineHFT before invoking OCR APIs. A non-null function table alone does not establish readiness: missing dltesseract5.dll or tessdata4 caused errors when OCR was requested. Handle initialization and recognition failures and guard cleanup.
If OCR remains enabled, language-data reduction requires separate validation for every supported language and recognition mode; selective trimming was not established by this investigation. Retain APDFL core dependencies and required font, CMap, Unicode, and color-profile resources. Validate the final product's retained workflows in an isolated deployment.
How to Get Additional Help
Contact Datalogics Tech Support
Fetch the answers quickly using our AI assistant Scout, available for any language. Scout also opens access to our MCP server, connecting APDFL with your favorite AI tools for faster and easier coding.