Extracting text and tables from scanned PDFs — 100% offline with WebAssembly

6 min read
pdf
ocr
tables
csv
privacy

Bank statements, invoices, medical forms and signed contracts are exactly the documents you should never upload to a random converter. They are also the documents that most often arrive as scans. Both tools below read them inside your browser tab.

Two kinds of PDF

A PDF made by a word processor stores real characters and their positions. A scanned PDF stores only a picture of each page. The PDF Text Extractor checks every page: when a page has a text layer it reads it instantly and losslessly; when a page has almost no text it renders the page at high resolution and runs OCR on it.

  • Embedded text is always preferred — it is exact and immediate
  • Scanned pages are handled automatically with Auto-OCR
  • OCR every page forces recognition when a PDF has a broken or garbled text layer
  • A page range such as 1-3, 5 keeps long documents fast

How OCR runs without a server

The OCR engine is Tesseract compiled to WebAssembly. It loads only when a page actually needs it, downloads the language data once, and then works on your machine. Each page is drawn onto a canvas and recognised; the canvas is cleared and the engine shut down when the job finishes, so memory does not pile up.

Tables from word positions

OCR and PDF text both come with coordinates. The extractor groups words into lines by their vertical position, splits each line where the horizontal gap is wide, and lines those cells up into columns that repeat across rows. The result appears in the Tables tab as a preview grid.

  • CSV for spreadsheets and databases
  • TSV to paste straight into Excel or Google Sheets
  • Markdown tables for READMEs and notes
  • JSON arrays that use the first row as keys

Then open the CSV in the CSV Viewer & Editor to rename columns, remove empty rows and sort.

Photos, receipts and batches

For images, use Image to Text. Drop a whole folder of receipts or book pages, click Extract all, and download every result as one file. Choose the Columns / table layout for tabular images. The tool lists words it is unsure about and offers one-click fixes for hyphenated line breaks and extra spaces.

Getting better results

  • Scan at 300 DPI or photograph straight-on in good light
  • Pick the correct OCR language before running
  • Keep Enhance image on for faded or low-contrast photos
  • Check confidence: below about 80% usually means a blurry or skewed source

Nothing in this workflow is uploaded, logged or stored. Close the tab and the document is gone.

Tools from this article

← All articles