DocxIntel home
DocxIntel, a product of BizfyLabs
DocxIntel, a product of BizfyLabs
by
BizfyLabs
  • Capabilities
    • Analyse

      Resolve layout, reading order, tables and handwriting

    • Identify

      Pull entities, fields and clauses with coordinates

    • Classify

      Sort document types and split multi-page packets

    • Map

      Link and reconcile entities across your estate

    • Modify

      Redact, mask and transform documents safely

    • Ask

      Query your documents and get cited answers

    • All six capabilities, one platform→
  • Deployment
  • Accuracy
  • Industries
    • Banking & Financial Services

      Statements, KYC files, and financial filings

    • Insurance

      Claims, policies, and underwriting documents

    • Government & Public Sector

      Records, correspondence, and regulatory filings

    • Healthcare

      Patient records, referrals, and lab reports

    • Legal & Compliance

      Contracts, filings, and case documentation

    • Energy & Utilities

      Engineering documents, contracts, and reports

    • Every regulated industry we serve→
  • Pricing
  • Docs
Book a Demo
Home
DocXIntel

Menu

    • All capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
    • All industries
    • Banking & Financial Services
    • Insurance
    • Government & Public Sector
    • Healthcare
    • Legal & Compliance
    • Energy & Utilities
    • Deployment models
    • Reference architectures
    • Sizing & throughput
    • What's in the box
    • Security posture
    • Documentation
    • Accuracy benchmark
    • Pricing
    • Proof of Value
    • Compare
    • About
    • FAQ
    • Contact Us
Capability — Analyse

Understand every document, before you extract a single field.

Analysis is the layer everything else depends on. DocxIntel resolves page structure, reading order, tables, stamps and handwriting across Arabic and English — running entirely on your hardware, with coordinates and confidence attached to every block it returns.

  • Runs air-gapped
  • Arabic-native reading order
  • Coordinates on every block
  • No per-page charge
Book a technical walkthrough
See the accuracy benchmark
DocxIntel analysing a scanned document, showing detected text blocks, tables and reading order
99%
Field-level accuracy
on the published benchmark set
AR + EN
Native script support
including mixed pages
0 bytes
Leaving your network
air-gapped installs
Unmetered
Pages analysed
fixed annual licence

Accuracy is measured field-by-field against a human-adjudicated ground truth, not character-by-character. The methodology, document mix and per-field breakdown are published in full on the accuracy page.

Under the hood

What analysis resolves before extraction begins

Most extraction errors are not extraction errors. They are structure errors that happened one step earlier — a merged table, a mis-ordered column, a stamp read as a character.

Page structure and reading order

Headers, body text, columns, footnotes, captions and marginalia are separated and ordered the way a human reads them, including right-to-left flow on Arabic pages.

  • Multi-column and mixed-orientation pages
  • RTL and bidirectional text runs
  • Header and footer suppression

Table structure, not just table text

Cell boundaries, merged cells, spanning headers and continuation rows across page breaks are reconstructed into an addressable grid.

  • Borderless and ruled tables
  • Multi-page table continuation
  • Merged and spanning headers

Handwriting and annotations

Handwritten field entries, marginal notes, ticked checkboxes and initials are recognised as content rather than discarded as noise.

  • Handwritten Arabic and English
  • Checkbox and radio state
  • Marginal annotations

Stamps, seals and signatures

Overlapping ink is separated from the printed text beneath it, so a stamp across an invoice total does not corrupt the total.

  • Stamp-over-text separation
  • Seal and signature detection
  • Wet-ink and digital marks

Image quality recovery

Skew, warp, glare, shadow, low contrast and generational fax loss are corrected before recognition rather than passed through to it.

  • Deskew, dewarp and perspective
  • Glare and shadow suppression
  • Multi-generation fax recovery

Provenance on every element

Each block, cell and token carries its page number, bounding box and confidence score, so any downstream value can be traced back to the pixels it came from.

  • Bounding-box coordinates
  • Per-element confidence
  • Deterministic element identifiers
Auditability

An answer you cannot trace is an answer you cannot sign off.

Analysis output is not a flat blob of text. Every element keeps its position on the page, which is what lets a reviewer, an internal auditor or a regulator see exactly where a number came from.

  • Click any extracted value and the source region highlights on the original page
  • Confidence scores are exposed per field, not averaged into a single document score
  • Element identifiers are stable across re-runs, so diffs between model versions are reviewable
  • The full analysis payload is available over your own API, not through a vendor portal

This is the difference between a tool that gives you data and a tool that survives an audit.

See field traceability in detail
An Arabic claim form with extracted fields highlighted and traced back to their source coordinates
Reference

Inputs, outputs and limits

What the analysis service accepts and returns in a standard deployment.

Accepted inputs

PDF
Native, scanned, and hybrid PDFs, including encrypted files where you supply the key
Images
JPEG, PNG, TIFF (single and multi-page), BMP, WEBP, HEIC
Office documents
DOCX, XLSX, PPTX, plus legacy DOC, XLS and PPT
Email and archives
EML, MSG with attachments, ZIP archives expanded in place
Fax and low quality
CCITT Group 3/4 fax TIFFs, including multi-generation copies

Outputs

Structured JSON
Blocks, tables, tokens, coordinates and confidence scores
Markdown
Reading-order text with table structure preserved, for retrieval pipelines
Annotated PDF
Original pages with detected regions overlaid for review and sign-off
Tabular export
Per-table CSV or XLSX with merged-cell resolution applied

Operational limits

Pages per document
No fixed ceiling; multi-thousand-page packets are split and processed in parallel
Scripts
Arabic and Latin scripts natively, including bidirectional pages
Throughput
Scales horizontally by adding GPU workers — see sizing and throughput
Retention
Governed entirely by your own storage policy; DocxIntel holds no copy

Format coverage in a specific deployment is confirmed in writing during the Proof of Value against your actual document mix, not against a generic list.

Positioning

Why on-premise analysis is a different product, not a deployment option

Cloud parsing APIs are good software. They are also priced per page, and every page has to reach them to be read.

Why on-premise analysis is a different product, not a deployment option
CriterionMetered parsing APIEnterprise BYOC tierDocxIntel
Where pages are processedVendor cloud tenantYour cloud tenantYour infrastructure, including bare metal
True air-gapped operationNot availableRequires cloud connectivitySupported by default
Cost of re-analysing 2M archived pagesCharged again per pageat published per-page ratesCharged again per pageNo incremental charge
Model weightsVendor-hostedVendor-controlled imagesShipped to you in the bundle
Self-hosting available onNo planTop enterprise tier onlyEvery licence
Arabic as a primary scriptBest-effortBest-effortFirst-class

Where pages are processed

Metered parsing API
Vendor cloud tenant
Enterprise BYOC tier
Your cloud tenant
DocxIntel
Your infrastructure, including bare metal

True air-gapped operation

Metered parsing API
Not available
Enterprise BYOC tier
Requires cloud connectivity
DocxIntel
Supported by default

Cost of re-analysing 2M archived pages

Metered parsing API
Charged again per pageat published per-page rates
Enterprise BYOC tier
Charged again per page
DocxIntel
No incremental charge

Model weights

Metered parsing API
Vendor-hosted
Enterprise BYOC tier
Vendor-controlled images
DocxIntel
Shipped to you in the bundle

Self-hosting available on

Metered parsing API
No plan
Enterprise BYOC tier
Top enterprise tier only
DocxIntel
Every licence

Arabic as a primary script

Metered parsing API
Best-effort
Enterprise BYOC tier
Best-effort
DocxIntel
First-class

Comparison reflects publicly documented behaviour of metered document-parsing APIs as of 2026, including credit-based per-page pricing and enterprise-gated self-hosting. Verify against current vendor documentation before making a decision — we would rather you checked.

Compare the options in detail
How it runs

From ingest to structured output, inside your perimeter

  1. 1
    Step 1

    Documents arrive from your own systems

    Watched folders, S3-compatible object storage, your DMS, an email intake mailbox, or a direct API call. Nothing is uploaded anywhere — the service reads from the storage you already control.

  2. 2
    Step 2

    Pages are normalised and quality-corrected

    Deskew, dewarp, denoise and contrast correction run first, so recognition sees the best available version of each page rather than the version that came out of the scanner.

  3. 3
    Step 3

    Structure and reading order are resolved

    Layout segmentation separates regions and establishes reading order, including right-to-left flow, before any text is transcribed. Tables are reconstructed as grids, not as loose text.

  4. 4
    Step 4

    Text, handwriting and marks are recognised

    Printed text, handwriting, stamps, seals and checkbox states are recognised by models running on your GPUs, each result carrying coordinates and a confidence score.

  5. 5
    Step 5

    Output lands where your pipeline expects it

    Structured JSON, markdown, annotated PDFs or tabular exports are written back to your storage or returned over your API, ready for identification, classification and routing.

Questions engineers ask about analysis

The technical objections that come up in every evaluation, answered without hedging.

No. The layout, OCR and table-structure models ship inside the deployment bundle and load from local storage. A fully air-gapped install performs identical analysis to a connected one — there is no fallback call to an external service, so there is nothing to firewall.

Arabic is a first-class script, not a post-launch add-on. DocxIntel handles right-to-left reading order, Arabic-Indic and Western numerals, diacritics, and pages that mix Arabic body text with English headers or Latin-script product codes. Reading order is resolved per text block, so a bilingual table does not collapse into one scrambled string.

Every text block and table cell carries a confidence score, and every value keeps the page number and bounding-box coordinates it came from. Low-confidence regions are flagged rather than silently guessed, so your review queue receives the 2% that need a human instead of the 100% that do not.

Yes. Deskew, dewarp, perspective correction, glare suppression and contrast normalisation run before recognition, so a handheld photo of a stamped invoice on a desk is treated as a normal input rather than an exception. This is the failure mode that sends most extraction tools to manual review.

It is not priced per page. Analysis is part of the fixed annual DocxIntel licence, sized by deployment footprint rather than volume. Re-analysing your entire back catalogue after a model update costs nothing extra, which is what makes historical reprocessing practical.

Keep reading

Identify→Pull named entities, fields and clauses out of analysed pages with coordinates attached.Accuracy benchmark→The document mix, the methodology and the per-field numbers behind the 99% claim.Sizing and throughput→GPU, CPU and storage requirements for the volume you actually process.What's in the box→Every container, model and dependency that ships in the deployment bundle.Deployment models→Air-gapped, private cloud, managed single-tenant or evaluation sandbox.Pricing→Why a fixed annual licence beats per-page metering once volume is real.

Start with a paid Proof of Value.

Thirty to forty-five days, fixed fee, deployed in your environment. The accuracy threshold and the conversion price are both agreed in writing before we begin.

Talk to an engineer
How the Proof of Value works
  • No per-page metering
  • Runs in your environment
  • Written accuracy threshold
DocxIntel Logo

A product of BizfyLabs

Document intelligence that never leaves your building. Analyse, identify, classify, map, modify and ask — inside your own infrastructure.

BizfyLabs on LinkedInDocxIntel documentationBizfyLabs

Product

  • Capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
  • Accuracy benchmark
  • Pricing
  • Proof of Value

Technical

  • Deployment models
  • Reference architectures
  • Sizing & throughput
  • What's in the box
  • Security posture
  • Model licences
  • Documentation
  • API reference

Solutions

  • All industries
  • Insurance & TPAs
  • Healthcare
  • Banking & finance
  • Government
  • Legal
  • Energy & logistics

Compare

  • Compare approaches
  • LlamaParse alternative
  • Docsumo alternative
  • On-premise document AI

Company

  • About DocxIntel
  • FAQ
  • Partners
  • BizfyLabs
  • Careers
  • Contact

© 2026 BizfyLabs FZC LLC. All rights reserved.

DocxIntel™ is a product of BizfyLabs FZC LLC.

  • Privacy Policy·
  • Terms of Service·
  • Data Processing Addendum·
  • Acceptable Use·
  • Model Licences·
  • Security·
  • Cookies

Registered in the United Arab Emirates. Delivery partner: Bizfy Solutions LLP, Indore, India.

DocxIntel