← All practical guidesPDF & documents

Extract text from a PDF: when to use OCR

Published October 3, 2026 · ToolSorcerer Team

A PDF can contain selectable text, pictures of text, or a mixture of both. Start by checking one page. That small check tells you whether to extract the existing text or recognize the printed words in an image.

Check what kind of PDF you have

Open your document in a PDF viewer and try to select a sentence. If you can copy meaningful words, begin with the PDF to Text Extractor. A scanned page can also have an existing OCR text layer; extraction uses that layer without recognizing the image again.

If selection only highlights an entire picture, or copying gives nothing useful, try the image workflow below. A blank extraction does not prove the page is a scan: unusual fonts or missing character mappings can also prevent useful text from being extracted.

Extract selectable text from selected pages

  1. Open one PDF in the extractor. Enter a password you know if the document asks for one.
  2. Choose the pages you need. For example, apply 1, 3-5 to request pages 1, 3, 4 and 5 in document order.
  3. Start with detected line breaks. Use continuous text per page if you need a single stream of text.
  4. Review each completed page, then copy it or download combined TXT, separate page files in a ZIP, or the JSON report.

Continuous text replaces detected line breaks with spaces. It does not reconstruct paragraphs or remove hyphenation. For a two-column article, compare the first paragraph of each column with the source preview before accepting the reading order.

Recognize text in a scanned page

Use PDF to Images to turn the needed page into an image, then open it in Image to Text / OCR. Choose the language used in the document, straighten rotated text, and crop unrelated material before recognition. Start with one clear page rather than an entire document.

For example, a scanned receipt may extract no text from the PDF tool. Render its page, crop to the receipt, run OCR, and compare the date, amounts and item names with the image. A plausible-looking result can still contain incorrect digits.

Check the result before using it

  • Compare names, numbers and punctuation with the original.
  • Check tables and columns manually; plain text does not preserve their visual layout.
  • Review right-to-left text and unusual fonts for missing or reordered characters.
  • Keep the source PDF so you can verify passages later.

These tools process the document in the browser. Production pages also load advertising; local processing does not mean there are no third-party scripts. See the privacy information before handling confidential documents.

Questions & answers

Why is the extracted text empty?

The page may be blank, scanned, or encoded with unsupported character mappings. Inspect the source, then try rendering the page as an image and using OCR if it contains printed text.

Will TXT preserve my PDF tables?

No. TXT contains plain text. Approximate spaces and line breaks can help, but columns, fonts and table structure need manual review.

Can I extract only a few pages?

Yes. Select page checkboxes or apply a page range before extraction. The output follows the source document order.

Does OCR always read numbers correctly?

No. Compare important numbers with the image. Blur, rotation, low contrast and unusual characters can all cause recognition errors.

Ready to try it?

Start with PDF to Text Extractor, or choose another practical guide.