OCR Language Support
Languages supported by in-browser OCR across Web, Studio, Studio Desktop, and the Embeddable Widget
Scanned and image-based PDF pages are read with OCR before redaction. This applies to PDF Redaction Online, On-Premise Studio, Studio Desktop, and the Embeddable Widget — they all run the same in-browser OCR engine. The API and SDK use a separate, server-side OCR engine and are not covered by this page.
English text — and any page that already has a text layer — is recognized automatically with no setup. For other languages, pick a language pack in Settings → OCR → OCR language: each pack pairs Latin script with one additional script family, so choose the one that matches your documents.
| Language pack | Covers |
|---|---|
| Latin / CJK (default) | Latin script, Chinese, Japanese |
| Latin / CJK (medium, more accurate) | Same coverage as above, using a larger and more accurate model |
| Latin | French, German, Spanish, Portuguese, Italian, Dutch, Polish, Turkish, Vietnamese, Romanian, Swedish, Danish, Norwegian, Finnish, Czech, Slovak, Hungarian, Croatian, and 30+ more Latin-script languages |
| Latin + Cyrillic | Russian, Ukrainian, Belarusian, English |
| Cyrillic only | Russian, Ukrainian, Belarusian, Serbian (Cyrillic), Bulgarian, Mongolian, Kazakh, Kyrgyz, Tajik, Macedonian, Tatar, Bashkir, and 20+ more Cyrillic/Caucasian-script languages |
| Latin + Hindi (Devanagari) | Hindi, Marathi, Nepali, Bhojpuri, Maithili, Sanskrit, Konkani, and other Devanagari-script languages |
| Latin + Arabic | Arabic, Persian, Urdu, Pashto, Uyghur, Sindhi, Kurdish, Balochi |
| Latin + Greek | Greek |
| Latin + Korean | Korean |
| Latin + Thai | Thai |
| Latin + Tamil | Tamil |
| Latin + Telugu | Telugu |
| English only | English (fastest, smallest download) |
Only one language pack runs per session, and downloaded models are cached in your browser after the first use so switching back is instant.