Guide
What Is OCR? How to OCR a Scanned PDF or Image
A scanned document or a photo of a page is just a picture: you can’t search it or copy the words. OCR reads the letters and turns them back into text. Here is when you need it, how to get a clean result, and what to check afterwards.
Published
What does OCR stand for?
OCR stands for optical character recognition: software that finds the letters in a picture of text, such as a scan, a photo or a screenshot, and turns them into characters you can select, search, copy and edit. The result is plain text, not a copy of the page’s layout.
Do you actually need OCR?
OCR is for text that exists only as a picture:
- a PDF made by a scanner or a phone scanning app;
- a photo of a letter, receipt, notice or book page;
- a screenshot of text you can’t select.
Many PDFs are not scans. If you can select a word in your PDF reader and copy it, the text is already there, and copying it is exact while OCR is only very good. ToolSolvo checks this for you: when a PDF already contains text, the OCR tool says so, and PDF to Word can get the original text instead.
How to extract text from an image
- Open Image to Text (OCR) and choose a JPG, PNG or WebP image (up to 20 MB). iPhone HEIC photos must be saved as JPG first.
- Choose the language: English, Urdu, or English and Urdu.
- Press Extract text. The first time, your browser downloads the OCR engine (about 4 MB) and the language data (about 3 MB for English, 1 MB for Urdu) from ToolSolvo; later runs usually come from your browser’s normal cache. A progress bar shows the progress, and Cancel stops the run.
- Read the result in the text box, correct any mistakes there, then use Copy text or download the .txt file. Your corrections are included in both.
The reading happens on your own device, using Tesseract (an open-source OCR engine) compiled for the browser. Your file is not uploaded.
How to OCR a PDF
A scanned PDF is a set of page images, so OCR reads it one page at a time. With ToolSolvo:
- Check whether you need OCR at all: try to select a word in your PDF reader. If you can, the text is already there; PDF to Word gets it exactly. The OCR tool also tells you when a PDF already contains text.
- Open Image to Text (OCR) and choose the PDF (up to 50 MB).
- Choose All pages (up to 20) or type a range such as 1-5. For a longer document, read it in parts, for example 1-20 and then 21-40.
- Choose the language and press Extract text. Each page is drawn at 300 DPI and read in turn; the progress bar shows which page is being read, and a page that takes longer than 90 seconds is stopped and reported.
- Check the result. When several pages are read, each page’s text starts after a line such as “--- Page 2 ---”, and pages where the engine’s confidence is low are pointed out. Correct mistakes in the box, then copy the text or download it as a .txt file.
The result is plain text, not a searchable PDF or a Word document. To edit it in Word, paste it into a new document. PDF to Word does not run OCR: it converts PDFs that already contain text and refuses a fully scanned PDF, pointing you to the OCR tool instead. To split a long PDF into smaller files first, use Split PDF.
How to get a better scan
Most OCR mistakes start with the picture. Before you blame the tool, check the input:
- Resolution: Tesseract’s documentation says it works best on images of at least 300 DPI. For a scanner, choose 300 DPI. For a phone photo, fill the frame with the page rather than photographing it from across a desk. ToolSolvo renders scanned PDF pages at 300 DPI automatically.
- Straight and flat: skewed lines hurt recognition badly. Place the page square to the scanner or camera, and flatten folds and curled book pages.
- Even light and contrast: avoid shadows, glare and dim rooms. Dark text on a plain, even, light background works best.
- Crop to the text: remove desk edges, fingers and other pages, but leave a small white margin around the text.
- One document per image: photograph pages one at a time rather than two side by side.
How to review the text for mistakes
OCR errors are often plausible-looking, so read with the original next to you:
- Numbers, dates and amounts first. A wrong digit in a phone number, account number or total is the costliest mistake and the hardest to spot.
- Look-alike characters: “rn” and “m”, “0” and “O”, “1”, “l” and “I”, “5” and “S”, and punctuation such as commas and full stops.
- Split or joined words, like “F ebruary” in our test, and words broken across lines.
- Line breaks: each line of the original becomes a line of text, so paragraphs may need joining back together.
- The tool also shows Tesseract’s own confidence for each page and warns when it is low. Treat that as a hint about where to look, not as an accuracy score.
What OCR keeps and what it loses
The result is plain text. It keeps the recognised characters line by line, with a line break for each line of the original and, when several pages are read, a separator line before each page. It does not preserve:
- layout, fonts, bold or colours;
- tables and columns, which come out as lines of text and may be read in the wrong order (for a table in a PDF that already has selectable text, PDF to Excel gives rows and columns instead; it cannot read scans);
- images, logos, signatures and stamps;
- form fields and checkboxes;
- handwriting, which is usually not recognised reliably.
Languages the tool is not set up for (it offers English and Urdu) will come out garbled. Urdu printed in Nastaliq, the usual style for books and newspapers, is often misread, as the example above shows.
Privacy and limits
Your file stays on your device, and nothing is saved in your browser’s storage. ToolSolvo processes up to 20 pages per run and scales very large images to about 12 megapixels. A single page that takes longer than 90 seconds is stopped and reported, so the tool never claims to have read text it didn’t. In our desktop tests, a 20-page run took 43 to 56 seconds and used roughly 350–470 MB of extra memory while it ran; phones, especially older ones, will be slower.
Sources
- Tesseract documentation, Improving the quality of the output — “Tesseract works best on images which have a DPI of at least 300 dpi”; skewed pages severely reduce line segmentation quality; noise and uneven backgrounds reduce accuracy.
- Tesseract.js on GitHub (README) — a WebAssembly port of Tesseract that runs in a web worker; Apache-2.0 licence.
- Tesseract.js documentation, Local installation — the worker, core and language files can be hosted locally instead of on a CDN, which is how ToolSolvo keeps OCR on its own site.