DocxIntel home
DocxIntel, a product of BizfyLabs
DocxIntel, a product of BizfyLabs
by
BizfyLabs
  • Capabilities
    • Analyse

      Resolve layout, reading order, tables and handwriting

    • Identify

      Pull entities, fields and clauses with coordinates

    • Classify

      Sort document types and split multi-page packets

    • Map

      Link and reconcile entities across your estate

    • Modify

      Redact, mask and transform documents safely

    • Ask

      Query your documents and get cited answers

    • All six capabilities, one platform→
  • Deployment
  • Accuracy
  • Industries
    • Banking & Financial Services

      Statements, KYC files, and financial filings

    • Insurance

      Claims, policies, and underwriting documents

    • Government & Public Sector

      Records, correspondence, and regulatory filings

    • Healthcare

      Patient records, referrals, and lab reports

    • Legal & Compliance

      Contracts, filings, and case documentation

    • Energy & Utilities

      Engineering documents, contracts, and reports

    • Every regulated industry we serve→
  • Pricing
  • Docs
Book a Demo
Home
DocXIntel

Menu

    • All capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
    • All industries
    • Banking & Financial Services
    • Insurance
    • Government & Public Sector
    • Healthcare
    • Legal & Compliance
    • Energy & Utilities
    • Deployment models
    • Reference architectures
    • Sizing & throughput
    • What's in the box
    • Security posture
    • Documentation
    • Accuracy benchmark
    • Pricing
    • Proof of Value
    • Compare
    • About
    • FAQ
    • Contact Us
Capability — Modify

Redaction that removes the data, not just the view of it.

Modification produces the documents you can safely release: truly redacted, selectively disclosed, watermarked and converted. Redacted regions are removed from the pixels and from the underlying content stream, and every operation writes a log a regulator can read.

  • Pixels and text layer removed
  • Redaction log on every operation
  • You hold the key, always
  • Runs air-gapped
Book a technical walkthrough
Read the security model
DocxIntel redacting personal identifiers from a document and producing a watermarked disclosure copy
Pixel-level
Redaction depth
content stream rewritten
Every region
Logged and hashed
page and bounding box
Your KMS
Key custody
no vendor recovery path
0 bytes
Leaving your network
air-gapped installs

Redaction correctness is verified during the Proof of Value against your own documents and your own disclosure policies, including an adversarial check that removed content cannot be recovered from the output file.

Under the hood

What document redaction and transformation covers

Most redaction incidents are not failures of judgement. They are failures of implementation — a rectangle drawn over text that was still sitting in the file.

True redaction

Regions are rasterised, the underlying glyphs are deleted from the content stream, and metadata, bookmarks, annotations and embedded attachments carrying the same content are stripped before the file is re-written.

  • No recoverable text under the mark
  • Metadata and attachments cleaned
  • Output hashed and recorded

PII and PHI masking

Personal and clinical identifiers are detected with coordinates by the identification layer, then removed, masked or partially masked according to policy — a last-four-digits pattern rather than an all-or-nothing choice.

  • Arabic and English names and addresses
  • Emirates ID, passport, IBAN, TRN
  • Clinical identifiers and diagnoses

Selective disclosure

One source document produces different minimal outputs for different recipients, each generated from a named policy rather than from an analyst deciding in the moment what to hide.

  • Recipient-specific policy profiles
  • Whole-page and whole-document exclusion
  • Reproducible from policy version

Watermarking and stamping

Visible watermarks, recipient identifiers, distribution notices and page-level stamps are applied on output, so a leaked copy can be traced back to the disclosure that produced it.

  • Per-recipient watermark text
  • Arabic and English stamp rendering
  • Header, footer and diagonal placement

Format conversion

Documents are converted between the formats your downstream systems and archives actually require, including flattening to image-only PDF where a text layer must not survive at all.

  • PDF, PDF/A, TIFF, image-only PDF
  • DOCX, XLSX and structured JSON
  • Deterministic, reproducible output

Re-assembly of packets

Documents split during classification are recombined into a defined pack order with a generated index, so a disclosure bundle arrives as one coherent file rather than nineteen attachments.

  • Ordered packs with generated index
  • Per-document inclusion rules
  • Consistent pagination and bates-style numbering
Hard documents

A stamped, scanned, bilingual contract is the normal case.

Redaction tools that depend on a clean selectable text layer fail on exactly the documents that carry the most sensitive data: photographed forms, wet-ink stamped contracts, faxed medical annexes and handwritten Arabic. DocxIntel redacts from the analysed page, so the presence of a text layer is irrelevant.

  • Handwritten identifiers are located by coordinate and removed at pixel level
  • Stamps and signatures overlapping text are handled without corrupting the printed content around them
  • Arabic right-to-left text runs are redacted as coherent regions, not as scrambled character spans
  • Multi-generation faxes and phone photos are normalised first, so detection does not silently miss a field
  • Where a text layer does exist, it is rewritten rather than covered, and the output is re-hashed

If a redaction pipeline only works on native digital PDFs, it does not work on the archive you actually hold.

See how pages are analysed
A stamped and signed contract with sensitive clauses redacted and the surrounding printed text intact
Worked example

Producing a third-party disclosure pack

A motor claim file has to go to an independent repairer. They need the damage assessment and nothing else.

  1. 1
    Step 1

    You select a disclosure policy, not a set of boxes

    A named policy defines which document classes are included, which data types are removed, which are partially masked, and what watermark the recipient receives. Policies are versioned and reviewable.

  2. 2
    Step 2

    Targets are detected with coordinates

    The identification layer locates every personal identifier, clinical reference and third-party name in the file, each with a page number, bounding box and confidence score. Nothing relies on a keyword search.

  3. 3
    Step 3

    Low-confidence detections escalate to a human

    A detection below threshold is not quietly dropped. It is queued with the region highlighted, because the failure mode that matters here is leaving data in, not taking too much out.

  4. 4
    Step 4

    Redaction is applied and the file is re-written

    Regions are rasterised, glyphs deleted from the content stream, metadata and attachments cleaned, the medical report excluded entirely, and the output watermarked with the repairer name and disclosure date.

  5. 5
    Step 5

    The output is verified before it is released

    An automated check re-extracts text and re-runs detection against the produced file. If any redacted value is still recoverable, the artefact is rejected rather than delivered.

  6. 6
    Step 6

    A redaction log is written to your storage

    Document, policy version, every redacted region, the triggering data type, the operator, timestamps and hashes of both source and output — append-only, in your own systems, ready for a regulator.

Reference

Redaction targets, modes and output formats

What can be removed, how the removal behaves, and what the service can produce.

Redaction targets

Identity data
Arabic and Latin personal names, Emirates ID, passport and visa numbers, dates of birth
Financial data
IBANs, account and card numbers, TRN, salary figures, transaction references
Contact data
Addresses, PO boxes, phone numbers, email addresses, plate and vehicle identifiers
Clinical data
Diagnoses, procedure codes, medication names, clinician and facility identifiers
Commercial data
Pricing, margins, named clauses, counterparty names and signature blocks
Manual targets
Arbitrary regions, whole pages, whole documents and named clause spans

Modes and key custody

Irreversible mode
Content is absent from the output; no recovery path exists for anyone, including us
Reversible mode
Removed regions stored encrypted beside the document, unwrapped only with your key
Key holder
Your KMS or HSM. DocxIntel holds no key and has no escrow of its own
Partial masking
Last-four, first-initial, year-only and format-preserving patterns per data type
Post-output verification
Automated re-extraction and re-detection against the produced file before release

Output formats and artefacts

PDF family
PDF, PDF/A-2b and PDF/A-3b for archive, plus flattened image-only PDF
Images
TIFF (single and multi-page), PNG and JPEG per page
Office and data
DOCX, XLSX and structured JSON with redactions applied
Disclosure packs
Ordered multi-document bundles with generated index and consistent pagination
Redaction log
Append-only JSON per operation with regions, policy version, actor and source and output hashes

Format coverage and policy behaviour for a specific deployment are confirmed in writing during the Proof of Value against your actual document mix and your own disclosure rules.

Positioning

True redaction versus a black box and versus a cloud redaction API

There are two common ways to get this wrong: hide the data instead of removing it, or remove it correctly on hardware belonging to someone else.

True redaction versus a black box and versus a cloud redaction API
CriterionAnnotation-style black boxCloud redaction APIDocxIntel
Is the hidden text recoverableYes — select, copy or strip the layerUsually no, if configured correctlyNo — content is absent from the output
Metadata and attachments cleanedNoVaries by product and settingYes, including bookmarks and embedded files
Works on scans and handwritingCosmetically onlyDepends on the OCR tier purchasedYes, redaction is coordinate-based
Arabic and bilingual documentsNot applicableBest-effortFirst-class
Where personal data is processedYour desktopVendor cloud tenantYour infrastructure, including bare metal
Regulator-ready redaction logNoneVendor-side audit trailAppend-only log in your own systems
Who can reverse a redactionAnyone with the fileDepends on vendor key handlingOnly a holder of your key, or nobody
Cost of redacting 2M archived pagesAnalyst time, at scaleCharged per pageat published per-page list ratesNo incremental charge

Is the hidden text recoverable

Annotation-style black box
Yes — select, copy or strip the layer
Cloud redaction API
Usually no, if configured correctly
DocxIntel
No — content is absent from the output

Metadata and attachments cleaned

Annotation-style black box
No
Cloud redaction API
Varies by product and setting
DocxIntel
Yes, including bookmarks and embedded files

Works on scans and handwriting

Annotation-style black box
Cosmetically only
Cloud redaction API
Depends on the OCR tier purchased
DocxIntel
Yes, redaction is coordinate-based

Arabic and bilingual documents

Annotation-style black box
Not applicable
Cloud redaction API
Best-effort
DocxIntel
First-class

Where personal data is processed

Annotation-style black box
Your desktop
Cloud redaction API
Vendor cloud tenant
DocxIntel
Your infrastructure, including bare metal

Regulator-ready redaction log

Annotation-style black box
None
Cloud redaction API
Vendor-side audit trail
DocxIntel
Append-only log in your own systems

Who can reverse a redaction

Annotation-style black box
Anyone with the file
Cloud redaction API
Depends on vendor key handling
DocxIntel
Only a holder of your key, or nobody

Cost of redacting 2M archived pages

Annotation-style black box
Analyst time, at scale
Cloud redaction API
Charged per pageat published per-page list rates
DocxIntel
No incremental charge

Comparison reflects publicly documented behaviour of metered document-processing and redaction products as of 2026, including per-page pricing and enterprise-gated self-hosting. Verify against current vendor documentation before making a decision.

Read the security model

Questions compliance teams ask about redaction

The questions a data protection officer asks before signing off a disclosure process.

A cosmetic redaction draws a black rectangle as an annotation and leaves the original text and image data underneath, recoverable by anyone who selects the region, removes the annotation layer or runs a text dump. DocxIntel rasterises the redacted region, deletes the underlying glyphs from the content stream, strips the corresponding entries from metadata, bookmarks and attachments, and re-writes the file so the removed content is not present in the output at all.

Only if you choose an irreversible-by-default policy or an escrowed reversible one, and only you hold the key either way. In irreversible mode the source content simply does not exist in the output. In reversible mode the removed regions are stored encrypted alongside the document, and the key lives in your own KMS or HSM. DocxIntel never holds a key and never has a recovery path of its own.

Automatically, using the same identification layer that powers extraction, so Emirates ID numbers, IBANs, passport numbers, Arabic and English personal names, addresses, phone numbers, dates of birth and clinical identifiers are detected with coordinates already attached. You then set policy per data type and per recipient, and anything below the confidence threshold is escalated for human confirmation rather than left in the document.

A redaction log. Every operation emits an append-only record listing the document, the policy and policy version applied, each redacted region by page and bounding box, the data type that triggered it, the confidence score, the operator or service account responsible and a cryptographic hash of both the source and the output file. The log is written to your own storage and can be shipped to your SIEM.

Both regimes expect you to disclose no more personal data than the purpose requires. Selective disclosure makes that operational: the same source document produces a different, minimal output for a repairer, a reinsurer, a regulator and a court, each generated from policy rather than from an analyst deciding what to hide. Because everything runs on your hardware, no personal data crosses a border to be redacted.

Yes, and scanned documents are the normal case. Redaction works from the analysed page, so a handwritten Emirates ID number on a photographed form or an Arabic name in a right-to-left text run is located by coordinates and removed at the pixel level. There is no dependence on a selectable text layer existing in the first place.

Keep reading

All capabilities→How analyse, identify, classify, map, modify and ask compose into one platform.Identify→The coordinate-bound detection of personal data that redaction policy acts on.Classify→Split packets so a medical report can be excluded from a pack as a whole document.Map→Find every document that mentions one person before you decide what to disclose.Security→Access control, key custody, audit logging and the controls behind data minimisation.Deployment models→Air-gapped, private cloud, managed single-tenant or evaluation sandbox.

Start with a paid Proof of Value.

Thirty to forty-five days, fixed fee, deployed in your environment. The accuracy threshold and the conversion price are both agreed in writing before we begin.

Talk to an engineer
How the Proof of Value works
  • No per-page metering
  • Runs in your environment
  • Written accuracy threshold
DocxIntel Logo

A product of BizfyLabs

Document intelligence that never leaves your building. Analyse, identify, classify, map, modify and ask — inside your own infrastructure.

BizfyLabs on LinkedInDocxIntel documentationBizfyLabs

Product

  • Capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
  • Accuracy benchmark
  • Pricing
  • Proof of Value

Technical

  • Deployment models
  • Reference architectures
  • Sizing & throughput
  • What's in the box
  • Security posture
  • Model licences
  • Documentation
  • API reference

Solutions

  • All industries
  • Insurance & TPAs
  • Healthcare
  • Banking & finance
  • Government
  • Legal
  • Energy & logistics

Compare

  • Compare approaches
  • LlamaParse alternative
  • Docsumo alternative
  • On-premise document AI

Company

  • About DocxIntel
  • FAQ
  • Partners
  • BizfyLabs
  • Careers
  • Contact

© 2026 BizfyLabs FZC LLC. All rights reserved.

DocxIntel™ is a product of BizfyLabs FZC LLC.

  • Privacy Policy·
  • Terms of Service·
  • Data Processing Addendum·
  • Acceptable Use·
  • Model Licences·
  • Security·
  • Cookies

Registered in the United Arab Emirates. Delivery partner: Bizfy Solutions LLP, Indore, India.

DocxIntel