Redact PII in PDFs from your own code. ScaleDP (Python) and ScaleDP-TS (JavaScript) combine OCR, AI entity detection and redaction in one pipeline, running on your hardware or entirely in the browser.
pip install scaledpnpm install @stabrise/scaledpAlso available as:
Online ToolStudio (Cloud)Studio (Self-Hosted)Studio DesktopCloud APISelf-Hosted APIn8n IntegrationWidgetMCP ServerObsidian PluginAutomatically detects and redacts PII like names, locations, emails, and dates.
Leverages advanced AI to understand context and make precise redactions.
Built-in OCR with Tesseract, EasyOCR, Surya or docTR detects and redacts text in scanned images and PDFs.
Use it as a plain Python library on one machine, or scale the same pipeline out on Apache Spark to redact thousands of documents.
Works with common languages: English, Spanish, German, Italian, Russian and more.
Works with Pandas and Spark DataFrames, so redaction plugs into existing data workflows and machine learning pipelines.
Install ScaleDP and compose a redaction pipeline in a few lines: load documents into a Pandas or Spark DataFrame, run OCR, and detect PII with NER.
pip install scaledpfrom scaledp import *
spark = ScaleDPSession()
pipeline = PipelineModel(stages=[
PdfDataToImage(),
TesseractOcr(),
Ner(model="d4data/biomedical-ner-all"),
])
result = pipeline.transform(df)Documents never leave the browser. OCR, NER and redaction all run client-side.
Models execute locally over WebAssembly or WebGPU, with no server round-trip.
PaddleOCR text recognition detects text in scanned images and PDFs.
GLiNER lets you define entity types like "person" or "email" at call time, with no retraining.
YOLO-based detectors locate signatures and faces for redaction.
Import only what you need (/pdf, /ocr, /ner, /detect) and chain stages with a simple Pipeline API.
Install @stabrise/scaledp and compose the same kind of pipeline entirely in the browser using onnxruntime-web.
npm install @stabrise/scaledpimport { Pipeline } from '@stabrise/scaledp'
import { PdfToImage } from '@stabrise/scaledp/pdf'
import { PaddleTextRecognizer } from '@stabrise/scaledp/ocr'
import { GlinerNer } from '@stabrise/scaledp/ner'
const pipeline = new Pipeline([
new PdfToImage({ resolution: 300 }),
new PaddleTextRecognizer({ preset: 'v6-small' }),
new GlinerNer({ labels: ['person', 'email', 'phone'] }),
])
const rows = await pipeline.transform(file)Yes. ScaleDP and ScaleDP-TS are open source under the GNU AGPL-3.0 license, so you can use, modify and self-host them for free. AGPL-3.0 requires that if you modify the code and make it available to others, including over a network, you share your changes under the same license.
Use ScaleDP for server-side and batch jobs in Python, on a single machine or scaled out on Apache Spark. Use ScaleDP-TS when documents must stay on the user's device: it runs OCR, PII detection and redaction in the browser with onnxruntime-web.
No. ScaleDP works as a regular Python library on a single machine with Pandas DataFrames. Apache Spark is optional and only needed when you want to distribute processing across a cluster.
Yes. Both libraries include OCR. ScaleDP supports Tesseract, EasyOCR, Surya and docTR, and ScaleDP-TS runs PaddleOCR in the browser, so text in scanned pages and images can be detected and redacted.
No. ScaleDP-TS runs entirely client-side: PDF rendering, OCR and entity recognition execute in the browser over WebAssembly or WebGPU. Only the model files are downloaded; your documents never leave the device.
Choose the API if you want a managed service you can call from any programming language without hosting models yourself, or if you would rather not take on AGPL-3.0 obligations in your own codebase.
Need help integrating ScaleDP into your pipeline? Talk to our ML and data engineers about your most complex redaction and document processing challenges.