DocxIntel home
DocxIntel, a product of BizfyLabs
DocxIntel, a product of BizfyLabs
by
BizfyLabs
  • Capabilities
    • Analyse

      Resolve layout, reading order, tables and handwriting

    • Identify

      Pull entities, fields and clauses with coordinates

    • Classify

      Sort document types and split multi-page packets

    • Map

      Link and reconcile entities across your estate

    • Modify

      Redact, mask and transform documents safely

    • Ask

      Query your documents and get cited answers

    • All six capabilities, one platform→
  • Deployment
  • Accuracy
  • Industries
    • Banking & Financial Services

      Statements, KYC files, and financial filings

    • Insurance

      Claims, policies, and underwriting documents

    • Government & Public Sector

      Records, correspondence, and regulatory filings

    • Healthcare

      Patient records, referrals, and lab reports

    • Legal & Compliance

      Contracts, filings, and case documentation

    • Energy & Utilities

      Engineering documents, contracts, and reports

    • Every regulated industry we serve→
  • Pricing
  • Docs
Book a Demo
Home
DocXIntel

Menu

    • All capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
    • All industries
    • Banking & Financial Services
    • Insurance
    • Government & Public Sector
    • Healthcare
    • Legal & Compliance
    • Energy & Utilities
    • Deployment models
    • Reference architectures
    • Sizing & throughput
    • What's in the box
    • Security posture
    • Documentation
    • Accuracy benchmark
    • Pricing
    • Proof of Value
    • Compare
    • About
    • FAQ
    • Contact Us
Capability — Map

One entity, forty documents, three spellings of the same name.

Mapping links the same person, company, policy, vehicle or asset across your whole document estate, reconciles the values those documents disagree on, and builds a relationship graph you can query from your own systems. Every link keeps the evidence that justified it.

  • Arabic transliteration variants resolved
  • Mismatches flagged, not overwritten
  • Audit trail on every link
  • Graph exposed over your own API
Book a technical walkthrough
See all six capabilities
DocxIntel linking the same entity across multiple documents into a relationship graph with evidence on each link
Every link
Carries its evidence
sources and coordinates
AR + EN
Name variant resolution
transliteration aware
0 bytes
Leaving your network
no external enrichment call
Unmetered
Archive reprocessing
fixed annual licence

Resolution thresholds are yours to set. The precision and recall achieved on your own entity population is measured during the Proof of Value, on your documents, against a human-adjudicated sample.

Under the hood

What cross-document entity mapping resolves

Extraction tells you what a document says. Mapping tells you whether forty documents are talking about the same thing — which is the question most operational decisions actually depend on.

Entity linking across documents

The claimant on a form, the holder on a policy schedule and the payee on an invoice are resolved to one entity, with each mention retained as a separate observation rather than merged away.

  • People, organisations, assets and accounts
  • Every mention keeps its own source
  • Confidence score on every link

Arabic name variant resolution

Mohammed, Muhammad, Mohamad and محمد are treated as candidate renderings of one name, then confirmed or rejected against corroborating identifiers rather than accepted on spelling alone.

  • Transliteration and diacritic variance
  • Honorifics and name-order differences
  • Arabic company names with legal suffixes

Reconciliation across record types

A claim is checked against the policy that should cover it and the invoice that should support it, field by field, with the source of each value retained on both sides.

  • Claim against policy against invoice
  • Coverage dates, limits and excesses
  • Totals against itemised line sums

Mismatch and duplicate detection

Conflicting values, near-duplicate submissions and resubmitted documents with one changed digit are surfaced as findings with both versions shown, not resolved silently.

  • Conflicting-value findings with both sources
  • Exact and near-duplicate clustering
  • Changed-field highlighting

A graph across the whole estate

Entities and the relationships between them form a graph spanning every document you hold, so you can ask which policies touch one company, or which claims share a repairer.

  • Typed entities and typed relationships
  • Traversal from any document or entity
  • Queryable from your own systems

An audit trail per link

Each link records its creator, timestamp, active rules and thresholds, source documents and the coordinates of the values that supported it, including human confirmations and rejections.

  • Evidence stored with the assertion
  • Human decisions recorded and attributed
  • Reproducible under an earlier rule set
Worked example

Reconciling a motor claim against a policy and an invoice

Three documents, produced by three organisations, describing one event. Nothing in them is formatted the same way.

  1. 1
    Step 1

    Fields arrive already traced to their source

    Identification has produced typed values with page numbers, bounding boxes and confidence scores from the claim form, the policy schedule and the repair invoice. Mapping never works from loose text.

  2. 2
    Step 2

    Candidate entities are generated

    Names, identifiers, plate numbers, IBANs and addresses are normalised into comparable forms — including Arabic-Indic numerals and three transliterations of one family name — and blocked into candidate sets.

  3. 3
    Step 3

    Links are asserted only on combined evidence

    A name match alone is not enough. A matching Emirates ID number, a shared date of birth and a consistent address push the combined score past your threshold; a name match with a conflicting identifier does not.

  4. 4
    Step 4

    Reconciliation compares what each document claims

    Coverage dates against the incident date, the vehicle on the schedule against the vehicle repaired, the invoice total against the policy limit and against the sum of its own line items.

  5. 5
    Step 5

    Disagreements become findings, not overwrites

    The invoice quotes a chassis number that differs from the schedule by two characters. That becomes a mismatch record naming both documents, both values and both coordinate sets, routed to the adjuster.

  6. 6
    Step 6

    The graph is written to your systems

    Entities, links, findings and evidence land in your own database and are available over your own API, ready for the case management system, the fraud team and the data warehouse.

Arabic in practice

Transliteration is where most entity resolution quietly fails.

One person in a UAE document estate is routinely spelled four ways in Latin script and once in Arabic script, across an Emirates ID, a bank statement, a policy schedule and a handwritten form. A resolution engine tuned on English names treats them as four customers, and every downstream count is wrong.

  • Arabic-script and Latin-script renderings of one name are linked as the same entity, not deduplicated by luck
  • Diacritics, hamza and alif variants, and doubled consonants are normalised before comparison
  • Honorifics, patronymics and reordered given and family names are handled as structure, not as noise
  • Arabic-Indic numerals inside identifiers are normalised so a matching Emirates ID is recognised as matching
  • Company names carrying LLC, FZ-LLC, PJSC or their Arabic equivalents resolve to one organisation

This is the difference between an entity count you can report and an entity count you have to apologise for.

See how fields are extracted
An Arabic claim form with the claimant name field traced to its coordinates and linked to matching records
Positioning

Entity mapping versus spreadsheets and versus a RAG index

The two things organisations reach for instead are manual reconciliation and a vector index. One does not scale, and the other has no idea what an entity is.

Entity mapping versus spreadsheets and versus a RAG index
CriterionManual reconciliation in spreadsheetsRAG index with no entity modelDocxIntel
What the system reasons aboutRows a person typed inText chunks and embeddingsTyped entities, links and evidence
Arabic name variantsCaught only if the reviewer noticesSimilar text, no identity assertionResolved and confirmed against identifiers
Documents comparable at onceWhatever fits in an analyst working dayWhatever fits in the context windowThe whole estate
Conflicting values between documentsFound late, if at allBoth retrieved, neither reconciledExplicit mismatch finding with both sources
Near-duplicate detectionDepends on sort order and luckHigh similarity, no verdictClustered with changed fields highlighted
Auditability of a decisionA cell with no provenanceA chunk, sometimes a pageRules, thresholds, sources and coordinates
Where the result livesA spreadsheet on a single laptopA vector store with no relationshipsA graph in your database, on your API
Where processing happensYour desksOften a hosted embedding and LLM APIYour infrastructure, air-gapped if required

What the system reasons about

Manual reconciliation in spreadsheets
Rows a person typed in
RAG index with no entity model
Text chunks and embeddings
DocxIntel
Typed entities, links and evidence

Arabic name variants

Manual reconciliation in spreadsheets
Caught only if the reviewer notices
RAG index with no entity model
Similar text, no identity assertion
DocxIntel
Resolved and confirmed against identifiers

Documents comparable at once

Manual reconciliation in spreadsheets
Whatever fits in an analyst working day
RAG index with no entity model
Whatever fits in the context window
DocxIntel
The whole estate

Conflicting values between documents

Manual reconciliation in spreadsheets
Found late, if at all
RAG index with no entity model
Both retrieved, neither reconciled
DocxIntel
Explicit mismatch finding with both sources

Near-duplicate detection

Manual reconciliation in spreadsheets
Depends on sort order and luck
RAG index with no entity model
High similarity, no verdict
DocxIntel
Clustered with changed fields highlighted

Auditability of a decision

Manual reconciliation in spreadsheets
A cell with no provenance
RAG index with no entity model
A chunk, sometimes a page
DocxIntel
Rules, thresholds, sources and coordinates

Where the result lives

Manual reconciliation in spreadsheets
A spreadsheet on a single laptop
RAG index with no entity model
A vector store with no relationships
DocxIntel
A graph in your database, on your API

Where processing happens

Manual reconciliation in spreadsheets
Your desks
RAG index with no entity model
Often a hosted embedding and LLM API
DocxIntel
Your infrastructure, air-gapped if required

Retrieval-augmented generation is genuinely useful, and DocxIntel ships it as the Ask capability. The point is narrower: a similarity index is not an entity model, and asking it to reconcile records is asking the wrong component. Verify entity-resolution claims against current vendor documentation before making a decision.

See cited question answering
Reference

Entity types, resolution signals and outputs

What the mapping service models, what it resolves on, and what it hands back to your systems.

Entity types

Person
Names in both scripts, Emirates ID, passport, date of birth, contact details
Organisation
Trade name, legal name, trade licence number, TRN, establishment details
Policy or contract
Policy and contract numbers, parties, coverage period, limits and excesses
Asset
Vehicles by plate and chassis, properties, equipment serial numbers
Transaction
Claims, invoices, payments, purchase orders and their reference numbers

Resolution signals

Strong identifiers
Emirates ID, passport number, IBAN, TRN, trade licence, chassis and plate numbers
Supporting attributes
Normalised names, dates of birth, addresses, phone numbers and email addresses
Script normalisation
Arabic and Latin transliteration mapping, diacritic and numeral normalisation
Thresholds
Auto-link, review and reject bands set per entity type by you

Outputs and integration

Graph store
Typed entities and relationships in a database you operate, inside your network
API access
REST and gRPC traversal, lookup and subscription endpoints on your own hostnames
Findings feed
Mismatches, duplicates and threshold breaches emitted as events to your workflow engine
Audit record
Per-link creator, timestamp, rule version, sources and supporting coordinates
Export
Bulk JSON and tabular export of entities, links and evidence for your warehouse

Entity model, resolution thresholds and integration points for a specific deployment are agreed in writing during the Proof of Value, against your own entity population rather than a generic schema.

Questions data teams ask about mapping

What gets asked once someone realises the duplicate rate in the archive is not what the report said.

Resolution combines several signals rather than trusting string similarity alone. Arabic transliteration variants, honorifics, name-order differences and abbreviated middle names are normalised, then weighed against corroborating identifiers such as an Emirates ID number, a date of birth, an IBAN or a shared address. A link is only asserted when the combined evidence clears the threshold you set, and the evidence is stored with the link.

It records the disagreement instead of picking a winner. A policy schedule saying one vehicle and a repair invoice saying another produces a mismatch record naming both source documents, both values and both sets of coordinates. Resolution is a business decision, so DocxIntel surfaces the conflict to whoever owns it rather than silently overwriting one value with the other.

Yes. Entities, links, mismatches and evidence are exposed over your own API and stored in a database you operate, so your case management system, data warehouse or investigation tooling can query the graph directly. There is no vendor portal that owns the graph and no export step to get your own relationships back.

Every link carries who or what created it, when, which resolution rules and thresholds were active, the source documents on both sides and the coordinates of the values that supported it. Human confirmations and rejections are recorded the same way. That means a reviewer can reconstruct exactly why two records were treated as one person, which is the question an auditor actually asks.

That is one of the most common first uses. Because reprocessing is unmetered, you can run resolution across the entire archive rather than a sample, and duplicates surface as clusters of documents describing one entity or one transaction. Near-duplicates — a rescanned invoice, a resubmitted claim with one changed digit — are flagged with the differing fields highlighted.

No. Resolution, reconciliation and graph storage all run inside your infrastructure on models shipped in the deployment bundle. There is no enrichment call to a third-party identity database and no fallback lookup, which is what makes the capability usable on an air-gapped network holding regulated customer data.

Keep reading

All capabilities→How analyse, identify, classify, map, modify and ask compose into one platform.Identify→The typed, coordinate-bound fields that entity resolution runs on.Classify→Split packets into real documents so links are made between documents, not page ranges.Ask→Cited natural-language answers over the estate, grounded in the same entity graph.Security→Access control, audit logging and key custody across the platform.Sizing and throughput→What resolving an entire archive actually costs in GPUs, CPU and storage.

Start with a paid Proof of Value.

Thirty to forty-five days, fixed fee, deployed in your environment. The accuracy threshold and the conversion price are both agreed in writing before we begin.

Talk to an engineer
How the Proof of Value works
  • No per-page metering
  • Runs in your environment
  • Written accuracy threshold
DocxIntel Logo

A product of BizfyLabs

Document intelligence that never leaves your building. Analyse, identify, classify, map, modify and ask — inside your own infrastructure.

BizfyLabs on LinkedInDocxIntel documentationBizfyLabs

Product

  • Capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
  • Accuracy benchmark
  • Pricing
  • Proof of Value

Technical

  • Deployment models
  • Reference architectures
  • Sizing & throughput
  • What's in the box
  • Security posture
  • Model licences
  • Documentation
  • API reference

Solutions

  • All industries
  • Insurance & TPAs
  • Healthcare
  • Banking & finance
  • Government
  • Legal
  • Energy & logistics

Compare

  • Compare approaches
  • LlamaParse alternative
  • Docsumo alternative
  • On-premise document AI

Company

  • About DocxIntel
  • FAQ
  • Partners
  • BizfyLabs
  • Careers
  • Contact

© 2026 BizfyLabs FZC LLC. All rights reserved.

DocxIntel™ is a product of BizfyLabs FZC LLC.

  • Privacy Policy·
  • Terms of Service·
  • Data Processing Addendum·
  • Acceptable Use·
  • Model Licences·
  • Security·
  • Cookies

Registered in the United Arab Emirates. Delivery partner: Bizfy Solutions LLP, Indore, India.

DocxIntel