PDF OCR (Text Recognition)
Turns a scanned PDF or a photo into a searchable PDF or plain text.
This tool runs on JavaScript. To use it, turn on JavaScript in your browser and reload the page.
SCANNED PAGE, SEARCHABLE TEXT.
PDF OCR reads the writing in a scanned PDF or in a photo of a document and turns it into text. Recognition runs on the RUPO server with the open-source Tesseract engine; the document is not sent to any other service. Each PDF page is drawn at 300 dpi and read one by one. With “Searchable PDF” the pages look exactly the same, with an invisible text layer added underneath: you can select and copy the text and search the document. “Plain text (TXT)” gives only the writing that was read; with several pages, the pages are separated by numbered lines.
“Document language” is the page language by default; if the document is in another language, pick it from the list. Turkish, English, German, French, Spanish, Italian, Portuguese, Japanese, Korean, Hindi and Arabic are supported; English words are also read in every language. The right language matters for letters such as é, ü, ñ and ç to come out correctly. For a PDF, the “Pages” field lets you read only some pages (for example 1, 3-5); at most 200 pages are processed at once. JPG, PNG, WebP, TIFF and BMP images can be uploaded too. Up to 8 files can be uploaded at once; all are processed with the same settings and download as a single ZIP. Handwriting and very blurry images are not read.
From start to finish
- Upload your document. A scanned PDF or a JPG, PNG, WebP, TIFF or BMP image; up to 8 files at once.
- Pick the output and language. Searchable PDF or plain text. If the document isn't in the page language, pick the right one from the “Document language” list.
- Run it and download. The result card shows the page and character count. Download; when you leave the tool, the copy on the server is deleted.
Frequently asked questions
Do accented letters come out right?
Yes. With the right document language, letters such as é, ü, ñ and ç are read correctly; on an English page, English is already the default. If the wrong language is picked, these letters can turn into plain ones; then pick the right language and try again.
What's the difference between a searchable PDF and plain text?
In a searchable PDF the page looks exactly as scanned; thanks to the invisible text layer, the writing can be selected, copied and searched. Plain text gives only the writing; formatting, images and table layout are not carried over. If you'll paste the text somewhere else, pick TXT; if you'll keep the document and search in it, pick searchable PDF.
What happens with a PDF that already contains text?
The pages are read again from the image, and a note on the result card says so. You don't need OCR to get the text of such a PDF; “PDF → Word” extracts it directly and faster.
Does my file stay on the server?
No. Your file reaches the server only to be processed and is deleted when the job is done; the output is deleted too once you download it and leave the tool. You don't need an account, a name or an email address, and your files are never sent to a third-party service. It is free to use: we set no hourly or daily quota on the number of jobs, we don't make you watch an ad before processing, and we add no watermark to the output.
Other PDF tools
RUPO Studio