What this tool does
This tool reads the printed or typed text inside an image — a screenshot, photo, scan, or receipt — and turns it into editable, copyable text. It uses optical character recognition (OCR), the same technique document scanners use, so you don't have to retype anything by hand.
How it works
The recognition runs on Tesseract, a long-established open-source OCR engine, compiled to WebAssembly so it runs directly in your browser. When you click Extract text, the engine and your chosen language data are downloaded once (then cached), the image is analysed locally, and the detected characters are grouped back into words and lines.
A confidence score gives a rough sense of how sure the engine is overall. Because no OCR is perfect, the result appears in an editable box so you can fix any stray characters before you copy or download it.
Tips for the best results
- Use a sharp, well-lit image where the text is in focus and not blurry.
- Higher contrast helps — dark text on a light background reads best.
- Straighten and crop to the text; skewed or rotated lines lower accuracy.
- Pick the matching language so accented or non-Latin characters are recognised.
Reading ID cards, forms, and busy backgrounds
ID cards, licences, and official forms are the hardest images to read, because their security patterns, holograms, logos, and photos all get misread as random extra characters. OCR works on a plain block of text, not a whole card at once.
The reliable approach is to turn on Read only a selected area and box one field at a time — just the name, then just the ID number, and so on. Reading a single line over a clean part of the card gives far cleaner results than trying to capture everything in one pass. For a card photo, set the image type to Photo or ID card.