Extract text from any PDF — including scanned documents — using OCR. Downloads as an editable Word document (.rtf).
Works on text PDFs and scanned image PDFs. Output is .rtf
How it works:
Text-based pages have their text extracted instantly via PDF.js.
Scanned or image pages are rasterized then read with OCR using Tesseract.js (English).
The tool automatically detects which method to use for each page.
OCR pages take longer — please be patient for large scanned files.