PDF Text Extractor that joins broken lines into paragraphs
Choose a PDF and the text of every page appears in the result box. Join the broken lines into paragraphs, then copy the pages you need or save them as TXT.
Drop one PDF here to pull out its text
The file is read in this browser and never uploaded to a server.
Up to 50 MiB · 200 pages · for PDFs that contain text (scanned images are not OCR’d)
This is a 4-page sample PDF we made. Pages 1–3 have a running header and page numbers; page 4 is just a drawing. Switch the line-break and header options to compare the results. Choose your own document above.
Reading 0 / 0 pages
Nothing extracted yet.
Choose a PDF and the text of every page will be read and shown here.
Text inside scanned images is not extracted because this tool does not run OCR. PDFs with an open password or copy restrictions are not processed. Text follows the order stored in the PDF, so multi-column layouts and tables may come out in a mixed order. TXT files are saved as UTF-8 with a BOM so Notepad and Excel open them correctly.
ready to use.
- Visible firstKeep the input and result positions clear.
- Results firstPut the main number up front and keep the process secondary.
- Less to askNo sign-up or extra information before using the tool.
How to pull text out of a PDF and clean it up
When you extract text from PDF files, you take the characters stored inside the file and turn them into plain, editable text. Copying straight from a PDF viewer usually adds a line break at the end of every visual line, so sentences arrive chopped into pieces. This tool reads the text page by page, joins the lines that belong to the same paragraph, and lets you copy only the pages you need or save them as a TXT file.
It works on PDFs that contain real text, such as documents exported from Word, Google Docs or Pages, and reports downloaded from the web. PDFs made by scanning paper contain only images; this tool does not run OCR, so those pages return no text and are listed separately as pages without text. Your file is read in this browser and is not sent to a server.
First, check that the PDF contains text
Two PDFs can look the same while storing very different things: one keeps characters as text, the other keeps each page as a picture. If you can drag across a sentence in a PDF viewer and it highlights word by word, the file contains text. If the whole page is selected as one image, or nothing can be selected, it is most likely a scanned PDF.
When you open a scanned PDF here, those pages are marked “no text”. Mixed files are common, for example an image-only cover followed by a text body. Check the page numbers listed under the result box and run only those pages through an OCR program.
From choosing a PDF to saving TXT
Choose or drop one PDF and the tool reads the text from the first page to the last once. A progress bar shows the page count while it reads, and then the text of all selected pages appears in the result box. Changing the line-break style or any option afterwards re-formats the result instantly without reading the file again.
Before trying a personal document you can press Open sample PDF to see what each option does. The sample has three pages with a running header and page numbers plus one drawing-only page, and it opens only when you press the button.
- Choose a PDF and check the file name, total pages and number of pages with text.
- Enter the pages to extract if you need fewer than all of them.
- Pick the line-break style, page markers, header and page-number removal, and space clean-up.
- Use View to read all selected pages or one page; a single page shows the original page beside the text.
- Take the result with Copy all, Copy page or Save TXT.
Join into paragraphs vs keep PDF lines
Most PDFs have no idea what a paragraph is. Each visible line is stored with its position, which is why a normal copy puts a line break after every line. Join into paragraphs merges the lines of one paragraph into continuous text and leaves one blank line between paragraphs.
For example, if a PDF shows “The data we collect includes your name and” on one line and “email address.” on the next, the result is “The data we collect includes your name and email address.” Keep PDF lines leaves the two lines as they are, which suits addresses, poems and lists where the line position itself carries meaning.
Paragraph boundaries come from the vertical spacing. When the gap between two lines is roughly 1.5 times the page’s usual line spacing, a new paragraph starts. Headings in a different font size, lines that jump back up because the text moved to the next column, and lines that start with a list marker such as •, 1. or (1) also begin a new paragraph.
How spaces are handled when lines are joined
English lines are joined with a single space. Chinese characters and Japanese kana are joined without a space, since those scripts do not separate words.
When a word is split with a hyphen at the end of a line, such as “exam-” followed by “ple”, the result is “exam-ple” with the hyphen kept. Deleting the hyphen would turn it into “example”, but the same rule would also break real hyphenated words like “well-known” or “COVID-19”. The tool keeps every character instead of guessing, so search for leftover line-end hyphens when you move text from justified academic papers.
Removing repeated headers and page numbers
Reports often repeat the document title at the top and a page number such as “Page 3 of 20” at the bottom of every page. Those lines end up scattered through the extracted text. With Remove repeated headers and page numbers turned on, the tool compares the first and last line of each page that has at least three lines (two lines at each edge when a page has five or more) and removes lines that repeat across pages. It never empties a page, so a page with a single chapter title stays intact.
Lines that differ only in their numbers count as the same line, so “Page 1 of 20” and “Page 2 of 20” are treated as repeats. A line is removed only when it appears on at least 60% of the compared pages and on at least three pages.
The option is off by default because a line you need, such as a table header repeated at the top of every page, can meet the same position rule. When the option is on, the number of removed lines is shown below the result; if the number looks wrong, turn it off and compare.
Extracting only the pages you need
Type 2,4-6 in Pages to extract and only pages 2, 4, 5 and 6 go into the result. Commas combine separate pages or ranges, and a hyphen marks a continuous range. Writing 6,2 still returns page 2 before page 6, and a page listed twice appears once.
The numbers count the PDF’s first page as 1. They can differ from the page numbers printed in the document: if a cover and a contents page come first, printed page 1 is page 3 here. Pick a single page in View and compare the original page preview to find the right number.
Page markers and pages without text
With Add page markers on, each page’s text starts with a line such as [Page 3]. Pages without text show as [Page 5 · no text], so you can see where the gaps are without leaving the result. Markers are handy when you need to cite page numbers.
With markers off, pages are separated by a blank line and pages without text are left out. A paragraph that continues onto the next page is still split at the page boundary, so check those spots when you move long text as one piece.
Saving TXT and opening it in other apps
Save TXT downloads the text of all selected pages under the original name plus _text, so report.pdf becomes report_text.txt. Even while you view a single page, the file contains the whole selection. To save one page only, enter just that page number in Pages to extract first.
The file is UTF-8 with a byte order mark (BOM). Excel and older versions of Windows Notepad may guess the wrong encoding for UTF-8 files without a BOM and show garbled accents or symbols; the BOM tells them the file is UTF-8. Line breaks are Windows-style (CR+LF), which Notepad, Word, and editors on macOS and Linux all read correctly.
When text comes out garbled or out of order
What you see in a PDF and what can be extracted are not always the same. Displaying a page only needs the shape of each glyph, but copying needs a map that says which character each glyph represents. PDFs missing that map produce symbols or unknown characters, and pages where many characters look like that are flagged as possibly garbled. Exporting the PDF again from the original app with fonts embedded often fixes it.
Accented letters stored as a base letter plus a separate accent are recombined into single characters, and ligatures such as “fi” or “fl” are expanded to normal letters, so searching the extracted text works as expected.
Text follows the order stored in the PDF. That usually matches reading order, but two-column papers, wide tables and captions inside figures can come out interleaved. Table rows become one line with the cells separated by spaces, so they will not paste into Excel as a table.
PDFs that cannot be processed and input limits
One PDF of up to 50 MiB and 200 pages can be processed at a time. MiB is a binary unit: 50 MiB is 52,428,800 bytes. Both limits apply separately, so a small file with 201 pages is still refused. Split longer documents into parts of 200 pages or fewer and extract them one after another.
PDFs protected with an open password are not processed because the tool has no password prompt. PDFs whose author restricted text copying are refused as well. PDFs that restrict only printing or editing while allowing copying are extracted normally.
- Empty (0-byte) or non-PDF files: export the document as PDF again from its original app. Renaming a file to .pdf does not work.
- Damaged PDFs: check whether another viewer opens the file, and if it does, use that viewer’s save-as-PDF option to create a clean copy.
- Several files at once: add one at a time. Adding a new file clears the previous document and results.
No upload, and what Reset clears
The contents of your PDF are not sent to or stored on a server for extraction. Reading the file, cleaning the text, drawing previews and creating the TXT all happen in the browser you are using, with no sign-up. The page itself still loads its layout, scripts and ads over the network, so this does not mean the web page makes no connections at all.
Reset, or adding another file, clears the document and results inside the tool. It does not delete your original PDF or any TXT you already downloaded. Copied text stays on the clipboard, so on a shared computer copy something else afterwards to overwrite it.
Five checks before you copy or save
Extracted text is the information stored in the PDF, not a copy of how the page looks. For important documents, run through these checks first.
- Compare the page and character counts under the result with what you expect; a very low count usually means some pages have no text.
- Look for notices about pages without text or possibly garbled text.
- For pages with columns or tables, pick the page in View and compare the order with the original.
- If header removal is on, check the number of removed lines and make sure nothing important disappeared.
- Double-check numbers, dates, amounts and names against the original, and open the saved TXT in the app where you will use it.
Frequently asked questions
I opened a scanned PDF and no text came out. Why?
A scanned PDF stores each page as an image, so there is no text to extract. This tool does not run OCR and marks those pages as having no text. Run the file through an OCR program to create a searchable PDF, then open that version here.
Copying from my PDF viewer breaks every line. Can I join them?
Yes. Choose Join into paragraphs and the lines of each paragraph are merged with single spaces, with one blank line left between paragraphs. Documents with irregular line spacing may split a paragraph in the wrong place, so read the result and fix those spots.
The saved TXT shows strange characters in Excel. What can I do?
The file is saved as UTF-8 with a BOM so most apps detect the encoding. If characters still look wrong, use the import option and choose UTF-8 (65001) as the file origin. If the file was re-saved by another app, its encoding may have changed at that point.
Will tables come out cell by cell?
No. Each table row comes out as one line with the cells separated by spaces. Cell borders and merged cells are not extracted, so pasting into Excel will not rebuild the table. Keep PDF lines makes it easier to compare rows with the original.
Does it keep the right order for two-column papers?
Usually, because most PDFs store the left column before the right one. Some apps store lines alternately, and then the order gets mixed. Select that page in View and compare it with the original preview to check.
Why can’t I extract text from a copy-protected PDF?
The author set a permission that does not allow copying text, and this tool does not bypass it. PDFs that only restrict printing or editing while allowing copying can be extracted. If you need the text, ask the author for a version without the copy restriction.
How do I extract a PDF with more than 200 pages?
Up to 200 pages are processed at once. Split the PDF into parts of 200 pages or fewer, extract each part, and paste the results together in order. Page markers then show the page numbers inside each part, not the original numbering.
I turned on header removal but the header is still there.
Lines are only treated as repeats when at least three pages can be compared and the header differs by nothing but numbers. Headers that are not stored as the first or last lines of each page are not removed either. In those cases, delete them from the result manually.
Can I edit the text in the result box?
The result box is read-only so that what you see always matches what is copied or saved. Copy the text or save the TXT, then edit it in Notepad, Word or any editor.
Words split with hyphens at line ends stay as “exam-ple”. Why?
Removing line-end hyphens automatically would also damage genuine hyphenated words such as “well-known” or “COVID-19”. The tool therefore keeps the hyphen and joins the two parts without a space. Use your editor’s find feature to review the remaining hyphens in justified text.