PDF Redaction

PDF Redaction SDK

PDF Redaction SDK

Leverage the PDF Redaction SDK for advanced PDF redaction with distributed processing, enabling high-performance automated data protection across large document sets.

ScaleDP

The PDF Redaction SDK is based on the open-source library ScaleDP — a Python library for processing documents using AI/ML pipelines, scaled with Apache Spark and Pandas DataFrames.

pip install scaledp
from scaledp import *

spark = ScaleDPSession()

pipeline = PipelineModel(stages=[
    PdfDataToImage(),
    TesseractOcr(),
    Ner(model="d4data/biomedical-ner-all"),
])

result = pipeline.transform(df)

Links: GitHub · PyPI · Docs

ScaleDP-TS

Prefer to keep documents client-side? ScaleDP-TS is our TypeScript sibling of ScaleDP that runs OCR, PII detection, and redaction pipelines entirely in the browser using onnxruntime-web — no document is uploaded to a server. It ships with PaddleOCR text recognition, GLiNER zero-shot NER, and YOLO detection (including signatures and faces), composed through the same pipeline model as the Python library.

npm install @stabrise/scaledp
import { Pipeline } from '@stabrise/scaledp'
import { PdfToImage } from '@stabrise/scaledp/pdf'
import { PaddleTextRecognizer } from '@stabrise/scaledp/ocr'
import { GlinerNer } from '@stabrise/scaledp/ner'

const pipeline = new Pipeline([
  new PdfToImage({ resolution: 300 }),
  new PaddleTextRecognizer({ preset: 'v6-small' }),
  new GlinerNer({ labels: ['person', 'email', 'phone'] }),
])

const rows = await pipeline.transform(file)

Links: GitHub · npm · Docs

Build PDF redaction into your platform

Start integrating our PDF redaction SDK. Connect with our lead ML/Data engineers today to discuss your most complex redaction and data processing challenges.

Contact Us

On this page