PDF Redaction

Also available as:

Online ToolStudio (Cloud)Studio (Self-Hosted)Studio DesktopCloud APISelf-Hosted APIn8n IntegrationWidgetMCP ServerObsidian Plugin

ScaleDP: PDF Redaction Library for Python

AI-Powered Redaction

Automatically detects and redacts PII like names, locations, emails, and dates.

Context-Aware Accuracy

Leverages advanced AI to understand context and make precise redactions.

Scanned Documents (OCR)

Built-in OCR with Tesseract, EasyOCR, Surya or docTR detects and redacts text in scanned images and PDFs.

Runs Locally or on Spark

Use it as a plain Python library on one machine, or scale the same pipeline out on Apache Spark to redact thousands of documents.

Multilingual

Works with common languages: English, Spanish, German, Italian, Russian and more.

Fits Your Data Stack

Works with Pandas and Spark DataFrames, so redaction plugs into existing data workflows and machine learning pipelines.

ScaleDP: Python Code Example

Install ScaleDP and compose a redaction pipeline in a few lines: load documents into a Pandas or Spark DataFrame, run OCR, and detect PII with NER.

pip install scaledp
from scaledp import *

spark = ScaleDPSession()

pipeline = PipelineModel(stages=[
    PdfDataToImage(),
    TesseractOcr(),
    Ner(model="d4data/biomedical-ner-all"),
])

result = pipeline.transform(df)

ScaleDP-TS: PDF Redaction Library for JavaScript

No Upload, No Server

Documents never leave the browser. OCR, NER and redaction all run client-side.

Runs on onnxruntime-web

Models execute locally over WebAssembly or WebGPU, with no server round-trip.

Browser-Based OCR

PaddleOCR text recognition detects text in scanned images and PDFs.

Zero-Shot NER

GLiNER lets you define entity types like "person" or "email" at call time, with no retraining.

Signature & Face Detection

YOLO-based detectors locate signatures and faces for redaction.

Modular & Composable

Import only what you need (/pdf, /ocr, /ner, /detect) and chain stages with a simple Pipeline API.

ScaleDP-TS: JavaScript Code Example

Install @stabrise/scaledp and compose the same kind of pipeline entirely in the browser using onnxruntime-web.

npm install @stabrise/scaledp
import { Pipeline } from '@stabrise/scaledp'
import { PdfToImage } from '@stabrise/scaledp/pdf'
import { PaddleTextRecognizer } from '@stabrise/scaledp/ocr'
import { GlinerNer } from '@stabrise/scaledp/ner'

const pipeline = new Pipeline([
  new PdfToImage({ resolution: 300 }),
  new PaddleTextRecognizer({ preset: 'v6-small' }),
  new GlinerNer({ labels: ['person', 'email', 'phone'] }),
])

const rows = await pipeline.transform(file)

Library or API: Which Do You Need?

Open-source library

  • You run it: on your servers, your Spark cluster, or in the user's browser
  • Python (ScaleDP) or JavaScript and TypeScript (ScaleDP-TS)
  • Free under AGPL-3.0
  • You choose and manage the OCR and NER models

PDF Redaction API & SDK

  • Managed REST API, also available self-hosted
  • Call it from any language over HTTP
  • Free tier, then pay per page
  • Production-tuned detection of PII, signatures, photos and QR codes
See the PDF redaction API & SDK

PDF Redaction Library FAQs

Yes. ScaleDP and ScaleDP-TS are open source under the GNU AGPL-3.0 license, so you can use, modify and self-host them for free. AGPL-3.0 requires that if you modify the code and make it available to others, including over a network, you share your changes under the same license.

Use ScaleDP for server-side and batch jobs in Python, on a single machine or scaled out on Apache Spark. Use ScaleDP-TS when documents must stay on the user's device: it runs OCR, PII detection and redaction in the browser with onnxruntime-web.

No. ScaleDP works as a regular Python library on a single machine with Pandas DataFrames. Apache Spark is optional and only needed when you want to distribute processing across a cluster.

Yes. Both libraries include OCR. ScaleDP supports Tesseract, EasyOCR, Surya and docTR, and ScaleDP-TS runs PaddleOCR in the browser, so text in scanned pages and images can be detected and redacted.

No. ScaleDP-TS runs entirely client-side: PDF rendering, OCR and entity recognition execute in the browser over WebAssembly or WebGPU. Only the model files are downloaded; your documents never leave the device.

Choose the API if you want a managed service you can call from any programming language without hosting models yourself, or if you would rather not take on AGPL-3.0 obligations in your own codebase.

Build PDF redaction into your platform

Need help integrating ScaleDP into your pipeline? Talk to our ML and data engineers about your most complex redaction and document processing challenges.

Contact UsCheck Your Redaction