Skip to main content

PDF OCR

Extract text from a scanned or image-only PDF online for free, directly in your browser. Supports English and Vietnamese, using on-device OCR — no upload or signup required.

Drop your file here

or choose a file

Supports PDF

Your files are processed locally in your browser and are not uploaded to our servers.

How to use PDF OCR

  1. 1

    Drop a scanned or image-only PDF onto the upload area, or click it to choose one.

  2. 2

    Choose which language(s) to recognize — English, Vietnamese, or both.

  3. 3

    Click "Extract text." Each page is rendered and OCR'd in your browser — this can take a while on longer documents.

  4. 4

    Copy the extracted text, or download it as a .txt file.

Common uses

  • Getting searchable, copyable text out of a scanned document or old paper form
  • Pulling text from a PDF made of photographed or screenshotted pages
  • Extracting a quote or passage from a scanned book or article

Supported formats and limitations

Accepts PDF, up to one file at a time.

  • OCR accuracy depends heavily on the scan's quality — a clean, high-resolution, non-skewed scan reads far better than a blurry photo or low-quality fax.
  • Only English and Vietnamese are supported. Other languages, handwriting, and stylized or decorative fonts are not reliably recognized.
  • This reads and rasterizes each page in your browser, so a long PDF can take a noticeable amount of time and memory — there's no server doing the work in the background.
  • Output is plain text only — original layout, columns, tables, and formatting are not preserved.

Frequently asked questions

No. Every page is rendered and OCR'd entirely in your browser using Tesseract.js — nothing is sent to a server.

English and Vietnamese, individually or together. Select whichever language(s) match your document before extracting.

OCR accuracy depends on scan quality — low resolution, skewed pages, or unusual fonts all reduce accuracy. Always review extracted text before relying on it.

No — the output is plain text in reading order, without columns, tables, or other formatting from the original page.

Yes, 40 MB.

Related tools