OCR for scanned PDFs and images
For PDFs that are actually scanned images, or for loose photos/images of a document (with no selectable text). We recognize the text with OCR. You can combine PDFs and images, in any order.
Languages to recognize
Only pick the languages that appear in your document — each one adds a several-MB download the first time it’s used.
Preprocessing before OCR
"Pure black and white" automatically calculates the exact cutoff point between ink and paper (Otsu’s method) — it can perform better than automatic contrast on very clean, single-color scans, but may lose detail on photos with uneven shadows or very fine text.
Limitations
The result looks the same as the original (unless you correct the tilt), but the text becomes selectable and searchable — you can also download just the recognized text as .txt. OCR gets slower the more pages/images and languages you pick, and it works best with sharp, straight scans — blurry or handwritten text, or a photo with strong shadows, may be recognized poorly or not at all. Automatic straightening fixes small tilts (up to about 9°); if your document is already straight or you want to keep it exactly as is, you can uncheck that box. About preprocessing: "Automatic contrast" works well in most cases; "Pure black and white (Otsu)" can give better results on very clean, even scans, but on photos with uneven shadows or highlighter-marked text it can lose information — if your scan is already sharp and in color, try "None".
