PDF to HTML Converter
Extract the text of a PDF into a simple HTML page.
Runs in your browser Your input never leaves your device.
Drop a PDF here, or .
About this tool
Choose a PDF and the tool reads it in your browser with PDF.js, pulls out the text of each page and builds a basic HTML file with one paragraph per page. The file is generated on your device and is not uploaded. This is text extraction, not a faithful conversion: headings, columns, tables, fonts, images and page layout are not kept, and text is joined with spaces in reading order as the PDF stores it. Scanned PDFs that contain only images produce little or no text because there is no OCR. You get a download button and a preview of the HTML code.
How to use it
- Drop a PDF onto the box or choose a file.
- Wait while the pages are read.
- Press Download HTML or copy the code shown.
- Open the file and edit it as needed.
Common problems
- The result is empty
- The PDF is probably a scan made of images. Run OCR on it first.
- Paragraphs run together
- Each page becomes a single paragraph; split it by hand afterward.
- Text order looks wrong
- Multi-column layouts are read in the order the PDF stores them.