OCR misses faint text after PDF recompression because low-contrast strokes can be downsampled, quantized, blurred, or merged into the page background.
A PDF viewer may make the recompressed page look acceptable through zoom interpolation, sharpening, and human contextual reading. OCR instead consumes a rendered raster at a selected resolution and converts local intensity patterns into characters. When faint ink occupies only a few pixels, small changes in compression or color processing can push those pixels below segmentation and recognition thresholds.
Recompression Can Rasterize and Downsample the Page
A PDF optimizer may resample embedded scans, reduce dots per inch, convert color spaces, or flatten transparency before applying a new codec. Thin strokes then cover fewer pixels, and averaging mixes their intensity with white background.
A broad review of document binarization describes how document binarization separates foreground from background under uneven illumination, degradation, and low contrast. Its evidence shows that faint-text recovery depends on information preserved before thresholding. This distinction remains visible during later household testing.
Repeated lossy saves compound the problem because each generation operates on prior artifacts. JPEG block transforms and quantization remove subtle high-frequency changes, while aggressive monochrome conversion can make a borderline gray stroke disappear entirely. The intermediate result must remain inspectable before automation follows.
OCR Rendering and Preprocessing Apply Additional Thresholds
The OCR engine or PDF library chooses a render DPI, antialiasing mode, color conversion, denoising, deskew, and binarization method. A page that remains readable at screen zoom may yield a lower-resolution or differently filtered OCR bitmap.
Research on OCR-aware binarization jointly optimizes binarization parameters for OCR, demonstrating that one thresholding choice does not serve every document. Low-contrast strokes can be classified as background even when a person infers them from surrounding words.
Adaptive thresholds can recover local faint text but may amplify paper texture and compression ringing into false characters. Global thresholds suppress noise more consistently yet fail where ink and background intensities overlap. That boundary should be measured separately under realistic operating conditions.
Human Vision and OCR Fail at Different Boundaries
People combine zoom, word shape, language expectation, and nearby sentences, while OCR requires enough local evidence for segmentation and classification. A plausible-looking page is therefore not proof that character pixels survived. The practical consequence appears when several sources compete for limited context.
An analysis of preprocessing effects on OCR connects preprocessing and OCR accuracy across degraded documents, reinforcing that resolution, contrast, noise, and thresholding interact rather than contributing independently. This dependency should remain explicit in the final interface.
The failure boundary is assigning every omission to recompression. The new OCR run may use another language pack, renderer, DPI, layout mode, or cached page image. Compare the exact rasters presented to the same engine.
Compare Source and Recompressed Pages at the OCR Input
Render original and recompressed pages at identical DPI and color settings. Measure pixel dimensions, codec, effective resolution, local stroke-to-background contrast, edge width, block artifacts, OCR confidence, character error rate, and missing regions using verified ground truth.
Use OCR inconsistency to check whether page-level inconsistency comes from layout or image quality. Repeat OCR with preprocessing fixed, then vary only render DPI, threshold method, and recompression quality. The result must therefore be checked against the original evidence.
Call recompression causal when the same engine succeeds on the original raster and fails where measurable stroke information fell. Preserve originals, avoid repeated lossy optimization, and choose a higher-quality or lossless path for documents whose faint marks carry legal or historical meaning.
Tech & AI HUB
More to Read

Why Do SMB File Changes Reach an Incremental Indexer in Bursts?
See how SMB write caching, leases, CHANGE_NOTIFY, buffer overflow, reconnect, and indexer batching reshape steady edits into bursty ingestion events.

Why Does Local AI Latency Oscillate With a Home Server Fan Curve?
See how heat, fan control, clock limits, sensor lag, and workload timing create periodic local AI latencyโand how to prove the relationship.

Why Does Media Indexing Heat a NAS More Than a Sequential Backup?
Learn why small-file I/O, codecs, thumbnails, metadata, databases, and AI inference create more component-wide heat than sequential copying.

