DocsBolt

Extract Text from PDF

Get the words out of a PDF as plain text — to paste into an email, feed into a script, or search through a long report without a PDF viewer.

.pdf · up to 200 MB per file

Use 1-based page numbers separated by commas.

  • PDF
  • up to 200 MB per file
  • Never uploaded

What Extract Text does

Sometimes the PDF is just in the way. You want the words in a spreadsheet cell, a chat message, a script, or a search — and copying page by page out of a viewer mangles the line breaks and drops half of it. This tool reads every page's text layer and writes it to a plain .txt file, with a marker between pages so you can find your place. There is no formatting because plain text has none; that is the point. If you need the headings and tables to survive, PDF to Word or PDF to Markdown keep them. If the file is a scan, there is no text layer to read and you will want OCR PDF first.

Formats and limits

Accepts
.pdf — a password-protected PDF needs its password, or Unlock PDF first
Produces
TXT
Files per run
One
Size limit
200 MB per file

How it works

  1. 01

    Upload the PDF

    Any PDF with selectable text.

  2. 02

    Choose pages

    Leave empty for the whole document.

  3. 03

    Download the .txt

    Pages are separated by a marker line.

Diagram: the text and picture blocks are lifted out of a page and delivered as a separate text file and image files.text.txtimg1.jpg

Limitations

  • A scanned PDF has no text layer, so the output is empty. Run OCR elsewhere first.
  • Reading order follows the PDF's internal order, which for multi-column layouts can interleave columns.
  • Hyphenated line breaks are kept as in the PDF; you may see words split across lines.

What happens to your file

Your file is never uploaded or stored.

Full detail in the Privacy Policy.

Questions

Why is my output empty or tiny?

The PDF has no text layer — it is a scan, i.e. pictures of pages. Run it through OCR PDF first, then extract the text from the result.

Does it keep the formatting?

No. Plain text has no bold, headings or tables. If you need those, use PDF to Word or PDF to Markdown.

What about text in two columns?

Columns are read in the order they were laid out in the file, which is usually left column then right. Occasionally a designer's file interleaves them; PDF to Word handles those better.

Why are some words joined together or split oddly?

The PDF stores text as positioned fragments, not as sentences, and a few generators leave out the spaces. The tool reconstructs them from position, which works for the vast majority of files; the odd exception comes from unusual typesetting.

Does it read text inside images?

No. Text that is part of a picture — a screenshot pasted into the PDF, a scanned page — is pixels, not characters. Run OCR PDF on the file first.

Can I extract just one page?

Yes. Enter the page number, or a range such as 5-9, in the page range box.

Tools people use next

More PDF tools

Convert, compress, merge, split, edit, secure and extract from PDFs.

Browse PDF tools