Extracting text and tables from scanned PDFs — 100% offline with WebAssembly
Bank statements, invoices, medical forms and signed contracts are exactly the documents you should never upload to a random converter. They are also the documents that most often arrive as scans. Both tools below read them inside your browser tab.
Two kinds of PDF
A PDF made by a word processor stores real characters and their positions. A scanned PDF stores only a picture of each page. The PDF Text Extractor checks every page: when a page has a text layer it reads it instantly and losslessly; when a page has almost no text it renders the page at high resolution and runs OCR on it.
- Embedded text is always preferred — it is exact and immediate
- Scanned pages are handled automatically with Auto-OCR
- OCR every page forces recognition when a PDF has a broken or garbled text layer
- A page range such as
1-3, 5keeps long documents fast
How OCR runs without a server
The OCR engine is Tesseract compiled to WebAssembly. It loads only when a page actually needs it, downloads the language data once, and then works on your machine. Each page is drawn onto a canvas and recognised; the canvas is cleared and the engine shut down when the job finishes, so memory does not pile up.
Tables from word positions
OCR and PDF text both come with coordinates. The extractor groups words into lines by their vertical position, splits each line where the horizontal gap is wide, and lines those cells up into columns that repeat across rows. The result appears in the Tables tab as a preview grid.
- CSV for spreadsheets and databases
- TSV to paste straight into Excel or Google Sheets
- Markdown tables for READMEs and notes
- JSON arrays that use the first row as keys
Then open the CSV in the CSV Viewer & Editor to rename columns, remove empty rows and sort.
Photos, receipts and batches
For images, use Image to Text. Drop a whole folder of receipts or book pages, click Extract all, and download every result as one file. Choose the Columns / table layout for tabular images. The tool lists words it is unsure about and offers one-click fixes for hyphenated line breaks and extra spaces.
Getting better results
- Scan at 300 DPI or photograph straight-on in good light
- Pick the correct OCR language before running
- Keep Enhance image on for faded or low-contrast photos
- Check confidence: below about 80% usually means a blurry or skewed source
Nothing in this workflow is uploaded, logged or stored. Close the tab and the document is gone.