DocxIntel home
DocxIntel, a product of BizfyLabs
DocxIntel, a product of BizfyLabs
by
BizfyLabs
  • Capabilities
    • Analyse

      Resolve layout, reading order, tables and handwriting

    • Identify

      Pull entities, fields and clauses with coordinates

    • Classify

      Sort document types and split multi-page packets

    • Map

      Link and reconcile entities across your estate

    • Modify

      Redact, mask and transform documents safely

    • Ask

      Query your documents and get cited answers

    • All six capabilities, one platform→
  • Deployment
  • Accuracy
  • Industries
    • Banking & Financial Services

      Statements, KYC files, and financial filings

    • Insurance

      Claims, policies, and underwriting documents

    • Government & Public Sector

      Records, correspondence, and regulatory filings

    • Healthcare

      Patient records, referrals, and lab reports

    • Legal & Compliance

      Contracts, filings, and case documentation

    • Energy & Utilities

      Engineering documents, contracts, and reports

    • Every regulated industry we serve→
  • Pricing
  • Docs
Book a Demo
Home
DocXIntel

Menu

    • All capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
    • All industries
    • Banking & Financial Services
    • Insurance
    • Government & Public Sector
    • Healthcare
    • Legal & Compliance
    • Energy & Utilities
    • Deployment models
    • Reference architectures
    • Sizing & throughput
    • What's in the box
    • Security posture
    • Documentation
    • Accuracy benchmark
    • Pricing
    • Proof of Value
    • Compare
    • About
    • FAQ
    • Contact Us
Capability — Classify

Know what every document is, and where it should go next.

Classification decides the type of each document, splits multi-page packets into the documents they actually contain, and routes each one to the team or queue that owns it. It runs on your hardware, against a taxonomy you define, with a confidence score on every decision.

  • Packet splitting with per-split confidence
  • Your taxonomy, not a vendor one
  • Arabic and bilingual documents
  • No per-page charge
Book a technical walkthrough
See all six capabilities
DocxIntel classifying a mixed document packet into typed documents and routing each to a downstream queue
99%
Field-level accuracy
on the published benchmark set
Auto-split
Multi-page packets
evidence-based boundaries
Yours
Document taxonomy
configured, not inherited
Unmetered
Documents classified
fixed annual licence

Accuracy is measured field-by-field against a human-adjudicated ground truth on a stated document mix. Classification performance against your own taxonomy is measured separately during the Proof of Value, on your documents.

Under the hood

What document classification has to get right

The hard part is almost never telling an invoice from a passport. It is the 40-page PDF that contains eleven documents, three of which are photographs of the same page.

Document-type classification

Each document is assigned a class from your taxonomy using layout geometry, template match, reading order and text together, rather than keyword matching that breaks on a reworded header.

  • Hierarchical classes and sub-classes
  • Template-level and semantic matching
  • Per-class confidence scores

Multi-page packet splitting

A single scanned PDF containing a whole submission is divided at real document boundaries, so downstream extraction sees eight documents instead of one confused one.

  • Layout and header discontinuity signals
  • Reference-number and orientation change
  • Confidence score on every split point

Routing to the right queue

Class plus extracted context decides the destination: a team, a workflow, a topic, a folder or an API callback. Routing rules are yours to write and version.

  • Rule-based routing on class and metadata
  • Priority and SLA tagging
  • Webhooks into your own workflow engine

Thresholds and human fallback

Accept, review and reject bands are set per class, because misrouting a brochure and misrouting a medical report are not the same mistake.

  • Per-class accept and review thresholds
  • Top-N candidate classes surfaced
  • Triage queue with evidence shown

Unseen document types

Anything that does not match is labelled unknown rather than forced into the nearest class. Recurring unknowns are clustered so you can add a class deliberately.

  • Explicit unknown class, not a silent guess
  • Clustering of repeated unknowns
  • One-click promotion to a new class

Arabic and bilingual packets

Arabic-only, English-only and bilingual versions of the same form resolve to the same class, and a packet that mixes scripts page by page splits correctly.

  • RTL layout as a classification signal
  • Script change as a boundary signal
  • Arabic class names and descriptions
Split and route

A 40-page claim packet, taken apart correctly

One PDF arrives from a broker. It contains everything, in no particular order, scanned at three different qualities. This is what happens to it.

  1. 1
    Step 1

    Every page is analysed before anything is decided

    Pages are deskewed, dewarped and contrast-corrected, then segmented into blocks and tables with reading order resolved. Classification never works from raw pixels or from a flat text dump.

  2. 2
    Step 2

    Boundary detection finds where documents change

    Layout discontinuity, header and footer change, reference-number change, orientation flip, script change and template change are scored at every page break. The packet is cut where the evidence agrees.

  3. 3
    Step 3

    Each fragment is classified against your taxonomy

    The 40 pages resolve into an Emirates ID copy, a driving licence, a police report, two repair invoices, a medical report, a claim form and a policy schedule — each with a class and a confidence score.

  4. 4
    Step 4

    Thresholds decide what a human sees

    Seven documents clear their accept thresholds. The eighth, a faxed medical annexe, lands at 0.71 against a 0.95 threshold and goes to triage with its two candidate classes and the evidence for each.

  5. 5
    Step 5

    Routing sends each document to its owner

    Invoices go to the motor assessment queue, the medical report to the clinical review queue under stricter access control, identity documents to KYC verification. Each destination is your rule, not a default.

  6. 6
    Step 6

    Downstream capabilities pick up clean inputs

    Because each fragment is now a known document type, identification runs the right schema against the right document, and mapping has clean entities to link across the packet.

The failure mode

An unsplit packet poisons everything downstream.

If a claim packet is treated as one document, the invoice schema runs against the police report, the medical report inherits the wrong access controls, and the extracted total belongs to no document in particular. Every error after that point is really this error, arriving late.

  • Extraction schemas are applied per document type, which requires knowing the type first
  • Access control follows document class, so a medical report inside a motor packet inherits the right restrictions
  • Retention policy is per document, not per upload — an ID copy and a repair invoice do not expire together
  • Retrieval and Q&A cite a document, so a citation into an unsplit 40-page blob is not a useful citation
  • Duplicate detection only works if the duplicate is a document rather than a page range

Classification is cheap insurance against every downstream capability being confidently wrong.

See how extraction uses the class
A multi-page document packet being separated into individually typed and routed documents
Reference

Taxonomy and routing configuration options

What you can define without a model change, and what the service returns for every decision it makes.

Taxonomy definition

Class hierarchy
Nested classes and sub-classes with no fixed depth limit
Class metadata
Name, description, Arabic label, business owner and retention class per entry
Examples
Reference documents attached per class to anchor template and layout matching
Versioning
Taxonomies are versioned; every classification records the version that produced it
Unknown handling
Explicit unknown class with clustering of recurring unmatched documents

Thresholds and controls

Per-class thresholds
Independent accept, review and reject bands for each class
Candidate output
Top-N candidate classes with scores, not a single opaque label
Split sensitivity
Boundary score threshold tunable per intake channel
Human override
Reviewer decisions recorded against document, class, user and timestamp

Routing targets

Rule inputs
Class, confidence, extracted fields, intake channel, language and page count
Destinations
Queues, folders, S3-compatible prefixes, database records, webhooks and message topics
Priority and SLA
Priority tags and SLA clocks assigned per class and per rule
Audit record
Class, confidence, taxonomy version, split points and routing decision, appended to your log

Taxonomy design is part of the Proof of Value. We build it against your real intake mix and measure classification performance on your documents, with the threshold agreed in writing up front.

Positioning

On-premise classification versus metered cloud classification

Cloud classifiers work. The question is what it costs to change your mind about your taxonomy, and whether a medical report has to leave the building to be recognised as one.

On-premise classification versus metered cloud classification
CriterionMetered cloud classifierIn-house rules and templatesDocxIntel
Where documents are classifiedVendor cloud tenantYour infrastructureYour infrastructure, including bare metal
Multi-page packet splittingAvailable on some productsUsually manual or page-count basedEvidence-based, with per-split confidence
Adding a new document typeVendor-side configuration or retrainingA new template and new brittle rulesA taxonomy entry with example documents
Reclassifying 2M archived documentsCharged again per pageat published per-page list ratesFeasible but rules rarely generaliseNo incremental charge
Behaviour on an unseen typeOften the nearest classNo match, no signalExplicit unknown, then clustered
Arabic and bilingual documentsBest-effortRarely handled at allFirst-class
True air-gapped operationNot availableYesSupported on every licence

Where documents are classified

Metered cloud classifier
Vendor cloud tenant
In-house rules and templates
Your infrastructure
DocxIntel
Your infrastructure, including bare metal

Multi-page packet splitting

Metered cloud classifier
Available on some products
In-house rules and templates
Usually manual or page-count based
DocxIntel
Evidence-based, with per-split confidence

Adding a new document type

Metered cloud classifier
Vendor-side configuration or retraining
In-house rules and templates
A new template and new brittle rules
DocxIntel
A taxonomy entry with example documents

Reclassifying 2M archived documents

Metered cloud classifier
Charged again per pageat published per-page list rates
In-house rules and templates
Feasible but rules rarely generalise
DocxIntel
No incremental charge

Behaviour on an unseen type

Metered cloud classifier
Often the nearest class
In-house rules and templates
No match, no signal
DocxIntel
Explicit unknown, then clustered

Arabic and bilingual documents

Metered cloud classifier
Best-effort
In-house rules and templates
Rarely handled at all
DocxIntel
First-class

True air-gapped operation

Metered cloud classifier
Not available
In-house rules and templates
Yes
DocxIntel
Supported on every licence

Comparison reflects publicly documented behaviour of metered document-processing products as of 2026, including per-page pricing and enterprise-gated self-hosting. Verify against current vendor documentation before making a decision.

Compare the options in detail

Questions operations teams ask about classification

What comes up when the intake queue is the bottleneck and nobody trusts an automatic label yet.

Boundaries are detected from evidence, not from page counts. Layout change, header and footer discontinuity, reference-number change, orientation shift, script change and form-template change are combined into a boundary score per page break. A 40-page claim packet becomes eight separate documents with a confidence score on each split.

It is labelled unknown with a confidence score and routed to a fallback queue rather than forced into the nearest existing class. DocxIntel also clusters recurring unknowns, so if the same unfamiliar form arrives two hundred times you get a prompt to add a class instead of two hundred separate exceptions.

You configure your own. Classes are defined as a hierarchy with names, descriptions, example documents, per-class confidence thresholds and the downstream queue each class routes to. There is no fixed vendor taxonomy you have to map your business onto, and changing the taxonomy does not require a retraining cycle.

Every class has its own accept and review threshold, because the cost of being wrong is not uniform. A misrouted marketing brochure is harmless; a misrouted medical report is a regulatory incident. Documents below the accept threshold go to a human triage queue with the top candidate classes and the evidence for each already shown.

Yes. Classification runs on the analysed page — layout, reading order, template geometry and text — so an Arabic-only form, an English-only form and a bilingual form of the same document type all resolve to the same class. Script is treated as a signal about the document, not as a barrier to reading it.

It costs nothing extra. Classification is part of the fixed annual licence, sized by deployment footprint rather than volume. That matters because taxonomies change: when you split one class into three, you can reclassify five years of documents overnight instead of writing a business case for it.

Keep reading

All capabilities→How analyse, identify, classify, map, modify and ask compose into one platform.Analyse→The structure and reading order that boundary detection and classification run on.Identify→Run the right extraction schema once you know what each document is.Modify→Redact and re-assemble the split documents into disclosure-safe packs.Sizing and throughput→GPU, CPU and storage requirements for the intake volume you actually process.Deployment models→Air-gapped, private cloud, managed single-tenant or evaluation sandbox.

Start with a paid Proof of Value.

Thirty to forty-five days, fixed fee, deployed in your environment. The accuracy threshold and the conversion price are both agreed in writing before we begin.

Talk to an engineer
How the Proof of Value works
  • No per-page metering
  • Runs in your environment
  • Written accuracy threshold
DocxIntel Logo

A product of BizfyLabs

Document intelligence that never leaves your building. Analyse, identify, classify, map, modify and ask — inside your own infrastructure.

BizfyLabs on LinkedInDocxIntel documentationBizfyLabs

Product

  • Capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
  • Accuracy benchmark
  • Pricing
  • Proof of Value

Technical

  • Deployment models
  • Reference architectures
  • Sizing & throughput
  • What's in the box
  • Security posture
  • Model licences
  • Documentation
  • API reference

Solutions

  • All industries
  • Insurance & TPAs
  • Healthcare
  • Banking & finance
  • Government
  • Legal
  • Energy & logistics

Compare

  • Compare approaches
  • LlamaParse alternative
  • Docsumo alternative
  • On-premise document AI

Company

  • About DocxIntel
  • FAQ
  • Partners
  • BizfyLabs
  • Careers
  • Contact

© 2026 BizfyLabs FZC LLC. All rights reserved.

DocxIntel™ is a product of BizfyLabs FZC LLC.

  • Privacy Policy·
  • Terms of Service·
  • Data Processing Addendum·
  • Acceptable Use·
  • Model Licences·
  • Security·
  • Cookies

Registered in the United Arab Emirates. Delivery partner: Bizfy Solutions LLP, Indore, India.

DocxIntel