DocxIntel home
DocxIntel, a product of BizfyLabs
DocxIntel, a product of BizfyLabs
by
BizfyLabs
  • Capabilities
    • Analyse

      Resolve layout, reading order, tables and handwriting

    • Identify

      Pull entities, fields and clauses with coordinates

    • Classify

      Sort document types and split multi-page packets

    • Map

      Link and reconcile entities across your estate

    • Modify

      Redact, mask and transform documents safely

    • Ask

      Query your documents and get cited answers

    • All six capabilities, one platform→
  • Deployment
  • Accuracy
  • Industries
    • Banking & Financial Services

      Statements, KYC files, and financial filings

    • Insurance

      Claims, policies, and underwriting documents

    • Government & Public Sector

      Records, correspondence, and regulatory filings

    • Healthcare

      Patient records, referrals, and lab reports

    • Legal & Compliance

      Contracts, filings, and case documentation

    • Energy & Utilities

      Engineering documents, contracts, and reports

    • Every regulated industry we serve→
  • Pricing
  • Docs
Book a Demo
Home
DocXIntel

Menu

    • All capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
    • All industries
    • Banking & Financial Services
    • Insurance
    • Government & Public Sector
    • Healthcare
    • Legal & Compliance
    • Energy & Utilities
    • Deployment models
    • Reference architectures
    • Sizing & throughput
    • What's in the box
    • Security posture
    • Documentation
    • Accuracy benchmark
    • Pricing
    • Proof of Value
    • Compare
    • About
    • FAQ
    • Contact Us
Capability — Identify

Extract the fields that matter, with the evidence still attached.

Identification turns a structured page into typed values: named entities, key-value fields, table line items and contract clauses. Every value returns with its page number, bounding-box coordinates and confidence score, so nothing you put in a downstream system is unverifiable.

  • Coordinates on every field
  • Custom schemas, no retraining
  • Arabic entity formats native
  • Runs air-gapped
Book a technical walkthrough
See the accuracy benchmark
DocxIntel extracting named entities and key-value fields from a document, each field boxed with a confidence score
99%
Field-level accuracy
on the published benchmark set
100%
Fields with coordinates
page and bounding box
No retraining
To add a field
schema configuration only
Unmetered
Fields extracted
fixed annual licence

Accuracy is measured field-by-field against a human-adjudicated ground truth, not character-by-character, and always against a stated document mix. The full methodology and per-field breakdown are on the accuracy page.

Extraction surface

What data extraction actually has to cover

A document is rarely a form. It is a form with a table in the middle, a stamp over the total, a handwritten correction in the margin and two pages of terms at the back.

Named entities

People, organisations, locations, dates, monetary amounts, identifiers and reference numbers are recognised wherever they appear, including inside free-flowing narrative text.

  • Arabic and Latin personal names
  • Company and trade names with legal suffixes
  • Dates in Hijri and Gregorian calendars

Key-value fields

Labelled fields are paired with their values using layout geometry rather than proximity guessing, so a value in the column to the left of its label on an Arabic form still binds correctly.

  • Printed and handwritten entries
  • Checkbox, tick and radio state
  • RTL label-to-value binding

Line items in tables

Invoice lines, statement rows, schedule entries and repair itemisations are extracted as typed rows from the reconstructed grid, not scraped from flattened text.

  • Merged and spanning header resolution
  • Rows continuing across page breaks
  • Row sums cross-checked against totals

Clause detection in contracts

Clauses are found by what they do — indemnity, termination, governing law, liability cap, renewal, confidentiality — rather than by matching a heading you hoped would be there.

  • Function-based, not heading-based
  • Full text plus page span returned
  • Deviation from your standard flagged

Custom field schemas

Define your own fields with names, types, validation rules and optional location hints. The extraction models are instructed against your schema at inference time, so there is no retraining loop.

  • Typed fields with validation rules
  • Per-field confidence thresholds
  • Versioned schemas, reviewable diffs

Provenance on every value

Each field returns the page it came from, the bounding box it occupies and a confidence score of its own. Nothing is averaged into a single document-level number.

  • Page and bounding-box coordinates
  • Per-field confidence
  • Stable identifiers across re-runs
Field-level traceability

Click the number, see the pixels it came from.

Field-level traceability is the difference between an extraction tool and a tool an auditor will accept. Every value DocxIntel returns keeps a pointer back to the exact region of the exact page it was read from, in the original document, months or years later.

  • Select any extracted field and the source region highlights on the original page
  • Arabic values are traced to their own bounding boxes, with RTL text runs preserved intact
  • Confidence is exposed per field, so a 0.72 remark does not hide behind a 0.99 document average
  • Field identifiers stay stable across re-runs, so a diff between two schema versions is reviewable
  • The whole traceability payload is available over your own API, not only inside a vendor portal

This is what an internal auditor, a regulator or a court asks for: not the value, but where the value came from.

See how accuracy is measured
An Arabic claim form with extracted fields highlighted and each value traced back to its source coordinates
Reference

Extractable field types and validation

The built-in types available before you define anything of your own. Custom types extend this list without a model change.

Identity and party fields

Person name
Arabic and Latin script, honorifics, multi-part family names, linked AR/EN renderings
Organisation name
Trade names, legal suffixes (LLC, FZ-LLC, PJSC), Arabic company names
Emirates ID
784-YYYY-NNNNNNN-C format with checksum validation
Passport and visa
Passport number, MRZ lines, issuing country, visa and residency numbers
Trade licence
UAE trade licence numbers, issuing authority, establishment card details

Financial fields

IBAN and account
AE and international IBANs with mod-97 checksum validation, account and SWIFT codes
Monetary amount
Currency detection, Arabic-Indic and Western numerals, bracketed negatives, written amounts
Tax registration
UAE TRN, VAT amounts and VAT-inclusive total checks
Line item row
Description, quantity, unit price, discount, tax and line total as typed columns

Dates, references and text

Date
Gregorian and Hijri, ambiguous format resolution against document context
Reference number
Policy, claim, invoice, purchase order, case and file numbers with pattern rules
Address
Arabic and English addresses, PO boxes, emirate and makani identifiers
Clause
Function-typed contract clauses with full text, page span and deviation flag
Free text
Remarks, narratives and descriptions with language tag

Validation and controls

Format rules
Regular expressions, checksum algorithms, enumerations and numeric ranges
Cross-field checks
Row sums against totals, date ordering, VAT arithmetic, required-if-present logic
Confidence thresholds
Set per field, with separate accept, review and reject bands
Output
JSON per document with values, types, coordinates, confidence and validation results

Field coverage for a specific deployment is confirmed in writing during the Proof of Value against your actual document mix and your actual schema, not against a generic list.

Human in the loop

Where a human belongs, and where they do not

The point of confidence scoring is not to eliminate review. It is to send your reviewers the two per cent that need judgement instead of the hundred per cent that do not.

  1. 1
    Step 1

    You define the schema and the thresholds

    Fields, types, validation rules and per-field accept and review thresholds are configured by your team. A policy number can demand near-certainty while a free-text remark is allowed to be approximate.

  2. 2
    Step 2

    Extraction runs against the analysed page

    Because analysis has already resolved structure and reading order, extraction works on a grid and a block tree rather than a wall of text. Each field is returned with coordinates and a confidence score.

  3. 3
    Step 3

    Validation and cross-field checks apply

    Checksums, format rules, date ordering and arithmetic checks run automatically. A total that does not equal the sum of its line items is flagged even when both values were read with high confidence.

  4. 4
    Step 4

    Only exceptions reach a reviewer

    Fields below threshold or failing validation are queued with the source region pre-highlighted. The reviewer sees the pixels and the proposed value side by side and confirms or corrects in one action.

  5. 5
    Step 5

    Corrections stay inside your perimeter

    Every correction is recorded against the field, the document and the reviewer, giving you an audit trail and your own labelled dataset. Nothing is sent anywhere to be trained on.

Positioning

On-premise data extraction versus metered extraction APIs

Per-page pricing is not just a cost question. It quietly decides which documents you can afford to look at, and how often you are willing to change your schema.

On-premise data extraction versus metered extraction APIs
CriterionMetered extraction APIEnterprise BYOC tierDocxIntel
Where fields are extractedVendor cloud tenantYour cloud tenantYour infrastructure, including bare metal
Adding a custom fieldPreset templates or vendor-assisted setupSame product, your tenantSchema change by your own team
Coordinates and per-field confidenceVaries by endpoint and presetVaries by endpoint and presetOn every field, always
Re-extracting 2M documents on a new schemaCharged again per pagepreset extraction is charged at a premium per-page rateCharged again per pageNo incremental charge
Arabic entity formatsBest-effortBest-effortFirst-class field types
Where reviewer corrections liveVendor platformVendor platformYour database, your dataset
True air-gapped operationNot availableRequires cloud connectivitySupported by default

Where fields are extracted

Metered extraction API
Vendor cloud tenant
Enterprise BYOC tier
Your cloud tenant
DocxIntel
Your infrastructure, including bare metal

Adding a custom field

Metered extraction API
Preset templates or vendor-assisted setup
Enterprise BYOC tier
Same product, your tenant
DocxIntel
Schema change by your own team

Coordinates and per-field confidence

Metered extraction API
Varies by endpoint and preset
Enterprise BYOC tier
Varies by endpoint and preset
DocxIntel
On every field, always

Re-extracting 2M documents on a new schema

Metered extraction API
Charged again per pagepreset extraction is charged at a premium per-page rate
Enterprise BYOC tier
Charged again per page
DocxIntel
No incremental charge

Arabic entity formats

Metered extraction API
Best-effort
Enterprise BYOC tier
Best-effort
DocxIntel
First-class field types

Where reviewer corrections live

Metered extraction API
Vendor platform
Enterprise BYOC tier
Vendor platform
DocxIntel
Your database, your dataset

True air-gapped operation

Metered extraction API
Not available
Enterprise BYOC tier
Requires cloud connectivity
DocxIntel
Supported by default

Comparison reflects publicly documented behaviour of metered extraction APIs as of 2026, including credit-based per-page pricing, preset extraction charged at higher rates and self-hosting gated to enterprise plans. Verify against current vendor documentation before making a decision.

Compare the options in detail

Questions engineers ask about extraction

The objections that decide an evaluation, answered without hedging.

No. Fields are defined in a schema — name, type, validation rule, optional hint about where it usually appears — and the extraction models are instructed against that schema at inference time. Adding a field is a configuration change your own team can make in an afternoon, not a labelling project and a retraining cycle.

Arabic entity handling is built into the field types, not layered on top of an English extractor. Arabic personal names with honorifics and multi-part family names, Arabic-script addresses, Emirates ID numbers, IBANs, trade licence numbers and Arabic-Indic numerals are all recognised in their native form. Where a value has both an Arabic and a Latin rendering on the same page, both are captured and linked.

It is flagged rather than guessed. Each field carries its own confidence score, and you set the threshold per field — a policy number might need 0.98 while a free-text remark tolerates 0.80. Fields below threshold land in a review queue with the source region already highlighted, so a human confirms in seconds instead of re-reading the document.

Yes, and this is where most extraction tools fail. Line items are read from the reconstructed table grid produced during analysis, so merged cells, spanning headers and rows that continue across a page break resolve correctly. Each line item is returned as a typed row with coordinates, and totals can be cross-checked against the sum of the rows.

Clauses are identified by function rather than by heading text, so a governing-law clause is found whether it is titled "Governing Law", "Applicable Law" or buried in a numbered miscellaneous section. Each detected clause returns its full text, its page span, its bounding boxes and a confidence score, which is what makes a contract review queue reviewable.

Neither. Identification is part of the fixed annual DocxIntel licence, sized by deployment footprint. Adding thirty fields to a schema or re-extracting two million archived documents against a new schema costs nothing extra, which is what makes iterating on a schema a normal engineering activity rather than a budget decision.

Keep reading

All capabilities→How analyse, identify, classify, map, modify and ask compose into one platform.Analyse→The structure, reading order and table reconstruction that extraction depends on.Classify→Decide what each document is and split packets before fields are extracted.Map→Link extracted entities across documents and reconcile the values that disagree.Accuracy benchmark→The document mix, the methodology and the per-field numbers behind the 99% claim.Security→How field-level audit trails, access control and key custody are handled.

Start with a paid Proof of Value.

Thirty to forty-five days, fixed fee, deployed in your environment. The accuracy threshold and the conversion price are both agreed in writing before we begin.

Talk to an engineer
How the Proof of Value works
  • No per-page metering
  • Runs in your environment
  • Written accuracy threshold
DocxIntel Logo

A product of BizfyLabs

Document intelligence that never leaves your building. Analyse, identify, classify, map, modify and ask — inside your own infrastructure.

BizfyLabs on LinkedInDocxIntel documentationBizfyLabs

Product

  • Capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
  • Accuracy benchmark
  • Pricing
  • Proof of Value

Technical

  • Deployment models
  • Reference architectures
  • Sizing & throughput
  • What's in the box
  • Security posture
  • Model licences
  • Documentation
  • API reference

Solutions

  • All industries
  • Insurance & TPAs
  • Healthcare
  • Banking & finance
  • Government
  • Legal
  • Energy & logistics

Compare

  • Compare approaches
  • LlamaParse alternative
  • Docsumo alternative
  • On-premise document AI

Company

  • About DocxIntel
  • FAQ
  • Partners
  • BizfyLabs
  • Careers
  • Contact

© 2026 BizfyLabs FZC LLC. All rights reserved.

DocxIntel™ is a product of BizfyLabs FZC LLC.

  • Privacy Policy·
  • Terms of Service·
  • Data Processing Addendum·
  • Acceptable Use·
  • Model Licences·
  • Security·
  • Cookies

Registered in the United Arab Emirates. Delivery partner: Bizfy Solutions LLP, Indore, India.

DocxIntel