How to use the PDF to Text Converter
- Choose the PDF you want to extract text from. Extraction starts page by page as soon as you select it.
- If you don't want separator lines such as "----- 1 -----" between pages, uncheck Add page separators. The result updates right away.
- Select the part you need in the results box and copy it, or click Copy all.
- Click Save as TXT to download a UTF-8 text file that opens in Notepad, TextEdit or any other text editor.
Why copying text from a PDF often goes wrong
When you drag to select text in a PDF viewer and copy it, the line breaks often come out jumbled. In two-column documents, sentences from the left and right columns can get mixed together, and odd spaces can show up between letters. This happens because a PDF only records which character to draw at which position. It doesn't store the structure of sentences or paragraphs.
This tool reads the coordinates of each text fragment in the PDF and groups fragments at the same height into one line. Where the gap between lines is larger than usual, it treats that as a new paragraph and adds a blank line. The result is usually much easier to read than text copied from a viewer.
When text extraction won't work on a scanned PDF
A PDF made with a scanner or a phone camera stores each page as a single image. The page looks like text to you, but the file has no text data in it, so the extraction comes back empty. To get text from these files you need OCR (optical character recognition) software, which recognizes letters in images. This tool doesn't support OCR.
To check whether a document is a scan, try selecting some text in your PDF viewer. If the whole page gets selected instead of individual words, the PDF is made of images.
Tips for using the extracted text
- When you quote from a report, paper or contract, you can paste the sentences directly instead of retyping them.
- If the extracted text has line breaks in the middle of sentences, use the Line Break Remover to join it back into paragraphs.
- To check length, paste the result into the Character Counter to see the character count and byte size.
- Tables are extracted as plain lines of text, without column breaks. If the table layout matters, get the original source file (for example the Excel spreadsheet or Word document) for accurate results.
- Text is extracted in page order, so long documents such as manuals or e-books come out in reading order and are ready to search or edit.