Scanned PDF to text

This tool converts scanned PDFs to editable text so you can easily copy, edit, and add content.

Upload your file

or drop, paste, or choose from cloud

Paste URL Try sample
Google Drive Dropbox OneDrive
100% Free No installation required Secure & private
Advertisement
Trusted by
eBay Booking.com University of Michigan Cornell University Columbia University
Rate this tool
0/5 · 0 votes

When a scanned PDF needs OCR, and when it does not

Turning a scanned PDF to text is really two different jobs, and the converter lets you choose which one runs. A PDF that already holds real characters only has to be read out. A PDF that holds pictures of pages has to be recognized first, because there is no text inside it yet.

You can tell them apart in a second. Open the file in any reader and try to select a line. If the words highlight, the PDF has a text layer, so pick Convert and the text comes out exactly as it was written. If nothing highlights, or the whole page highlights as one block, it is a scan and it needs optical character recognition.

That single choice explains most of the disappointing results people get elsewhere. A converter that runs one path for every upload returns clean text for some files and an empty page for others, with no explanation either way. Picking the engine yourself takes the guesswork out.

How to copy text from scanned PDF documents

A scan will not let you drag your cursor across a sentence, because there is nothing there to select. Extracting the text from a scanned PDF solves that: upload the file, run the conversion, and you get a plain text file you can copy from, paste into an email, or search with the find command in any editor.

The output is a .txt file, so it opens on any computer and on any phone without special software. It carries the words and the line breaks and leaves the layout behind, which is the right trade when what you want is the content rather than a copy of the page.

Several scans can go in at once. Each PDF is converted on its own and comes back as its own text file, with the same engine and the same language selection applied to all of them.

What is a TXT file?

The .txt file extension is used for plain text files. These files contain lines of text and can be opened in many text editors on different platforms and devices. The text does not contain any formatting.

That is why it is the safest place to put text recovered from a scan. Nothing about the file can go stale, nothing needs a licence to open, and the words stay readable long after the program that produced the original has been replaced.

File extension: .txt, MIME type: text/plain

TXT on Wikipedia

Which OCR engine to choose

There are 5 ways to read a scanned PDF here, and the settings panel lets you switch between them. Pick the one that matches the file in front of you.

OCR engine Choose it when
Convert The PDF already has selectable text. No recognition runs, so the words come out exactly as they were written.
Standard OCR The scan is clean, straight and sharp. This is the quickest route from a scanned PDF to text.
Advanced AI-OCR The scan is faint, low resolution or slightly crooked, and the standard engine drops characters.
Advanced AI-OCR+ The page has shadows, uneven lighting or a curved spine, which is common when a book is photographed.
Photo OCR The text was photographed rather than scanned, for example a sign, a label or a screen.

The converter recognizes 124 source languages, including simplified and traditional Chinese with vertical layouts and the Cyrillic forms of Azerbaijani, Serbian and Uzbek. Select every language that appears in your file.

Common reasons to convert a scanned PDF to text

The words in a scan are locked inside an image. These are the situations where getting them out saves the most work.

Paperwork you need to search

Invoices, contracts, receipts and letters usually arrive as scans. With the text extracted you can look up an amount, a date or a name instead of reading every page.

Study and research material

Library scans, journal articles and book pages become quotable. Extract the text once and paste the passages you need into your notes instead of retyping them.

Records and archives

Older files are often kept as long, multi-page scans. Upload several of them at once and the converter works through the batch in one run.

Your data, protected

The papers you scan and the documents you convert can be personal. Here is exactly how yours are handled.

TLS encryption

Every upload and download runs over an encrypted connection, so your files cannot be read in transit.

No human access

Text recognition is fully automated. Nobody reads your scans or documents, and nothing is shared with third parties.

ISO 27001 data centres

Processing happens in certified data centres inside the European Union.

How to convert a scanned PDF to text

1

Upload

Upload your scanned PDF.

2

Choose

Pick the OCR engine that fits your file. Use Convert if the text in the PDF is already selectable.

3

Adjust

Choose the language of your PDF for better results (optional).

4

Convert

Start the conversion and wait until your download is ready.

Scanned PDF to text FAQ

Is this scanned PDF to text converter free?

Yes. Converting a scanned PDF to text is free and needs no account. Upload the file, choose an engine and download the result. Files up to 100 MB are covered by the free limit, and a plan only matters for bigger files and longer batches.

Can I OCR a PDF for free?

Yes. Standard OCR is the free engine and it handles ordinary scans without a sign-up. The 3 advanced AI engines are the upgrade for pages the standard engine cannot read, such as faint copies or photos taken in poor light.

Why does my scanned PDF require OCR to extract text?

Because the file holds pictures of pages instead of characters. A scanner photographs the paper, so every word is a pattern of pixels. Optical character recognition reads those pixels and writes real characters, which is what makes the text selectable, searchable and editable.

How do I copy text from a scanned PDF?

Upload the PDF above, choose an OCR engine and start the conversion. The result is a plain text file you can open in any editor and copy from. Copying straight out of a PDF reader only works once the file has a text layer.

Which languages can the text recognition read?

It recognizes 124 source languages, including Chinese in both simplified and traditional form with vertical layouts, and the Cyrillic variants of Azerbaijani, Serbian and Uzbek. Select every language that appears in the file, because a mixed page reads best when all of them are listed.

Can I convert a scanned PDF to Word?

Not from this page, which writes a plain text file. It keeps the words and drops the layout. For a DOC or DOCX with formatting, use the Convert to Word app, which runs the same text recognition and writes a Word document.

Why did the extracted text come out wrong?

Recognition quality follows scan quality. A crooked page, a low resolution, a shadow across the paper or a photo taken by hand all cost accuracy. Try Advanced AI-OCR for imperfect scans and Advanced AI-OCR+ for poor lighting, and rescan at 300 dpi where you can. A password-protected or damaged source file will not convert at all.

Are my scanned documents private?

Yes. Uploads and downloads run over an encrypted connection, processing happens in certified data centres inside the European Union, and the whole pipeline is automated, so nobody reads your scans. Your files are never used to train AI models and you keep the rights to them.

More about text recognition

An open book on a yellow desk, framed by scanner brackets, with the article title Save Valuable Text With OCR printed across the page
How To

Save Important Text With OCR

Digitize important text from scanned documents and images with OCR (Optical Character Recognition). Extract text from images easily, online.

Read article