PDF Redaction

OCR Language Support

Languages supported by in-browser OCR across Web, Studio, Studio Desktop, and the Embeddable Widget

Scanned and image-based PDF pages are read with OCR before redaction. This applies to PDF Redaction Online, On-Premise Studio, Studio Desktop, and the Embeddable Widget — they all run the same in-browser OCR engine. The API and SDK use a separate, server-side OCR engine and are not covered by this page.

English text — and any page that already has a text layer — is recognized automatically with no setup. For other languages, pick a language pack in Settings → OCR → OCR language: each pack pairs Latin script with one additional script family, so choose the one that matches your documents.

Language packCovers
Latin / CJK (default)Latin script, Chinese, Japanese
Latin / CJK (medium, more accurate)Same coverage as above, using a larger and more accurate model
LatinFrench, German, Spanish, Portuguese, Italian, Dutch, Polish, Turkish, Vietnamese, Romanian, Swedish, Danish, Norwegian, Finnish, Czech, Slovak, Hungarian, Croatian, and 30+ more Latin-script languages
Latin + CyrillicRussian, Ukrainian, Belarusian, English
Cyrillic onlyRussian, Ukrainian, Belarusian, Serbian (Cyrillic), Bulgarian, Mongolian, Kazakh, Kyrgyz, Tajik, Macedonian, Tatar, Bashkir, and 20+ more Cyrillic/Caucasian-script languages
Latin + Hindi (Devanagari)Hindi, Marathi, Nepali, Bhojpuri, Maithili, Sanskrit, Konkani, and other Devanagari-script languages
Latin + ArabicArabic, Persian, Urdu, Pashto, Uyghur, Sindhi, Kurdish, Balochi
Latin + GreekGreek
Latin + KoreanKorean
Latin + ThaiThai
Latin + TamilTamil
Latin + TeluguTelugu
English onlyEnglish (fastest, smallest download)

Only one language pack runs per session, and downloaded models are cached in your browser after the first use so switching back is instant.