DocxIntel home
DocxIntel, a product of BizfyLabs
DocxIntel, a product of BizfyLabs
by
BizfyLabs
  • Capabilities
    • Analyse

      Resolve layout, reading order, tables and handwriting

    • Identify

      Pull entities, fields and clauses with coordinates

    • Classify

      Sort document types and split multi-page packets

    • Map

      Link and reconcile entities across your estate

    • Modify

      Redact, mask and transform documents safely

    • Ask

      Query your documents and get cited answers

    • All six capabilities, one platform→
  • Deployment
  • Accuracy
  • Industries
    • Banking & Financial Services

      Statements, KYC files, and financial filings

    • Insurance

      Claims, policies, and underwriting documents

    • Government & Public Sector

      Records, correspondence, and regulatory filings

    • Healthcare

      Patient records, referrals, and lab reports

    • Legal & Compliance

      Contracts, filings, and case documentation

    • Energy & Utilities

      Engineering documents, contracts, and reports

    • Every regulated industry we serve→
  • Pricing
  • Docs
Book a Demo
Home
DocXIntel

Menu

    • All capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
    • All industries
    • Banking & Financial Services
    • Insurance
    • Government & Public Sector
    • Healthcare
    • Legal & Compliance
    • Energy & Utilities
    • Deployment models
    • Reference architectures
    • Sizing & throughput
    • What's in the box
    • Security posture
    • Documentation
    • Accuracy benchmark
    • Pricing
    • Proof of Value
    • Compare
    • About
    • FAQ
    • Contact Us
Frequently asked questions

Every question a regulated buyer asks about on-premise document AI

Thirty-plus answers grouped by theme, written to be quotable in an internal review rather than skimmed in a sales meeting. Where the honest answer is uncomfortable — a certification we do not hold, a place where another product fits better — it is written that way.

  • No training on your data
  • Air-gapped by default
  • Fixed annual licence
  • Verifiable in your own logs
Ask us something else
Read the documentation
DocxIntel processing a stream of documents entirely inside a customer network
Never
Your data used for training
no mechanism, plus a clause
0 bytes
Leaving your network
air-gapped installs
99%
Field-level accuracy
on the published benchmark set
Unmetered
Pages processed
fixed annual licence

Accuracy is measured field by field against a human-adjudicated ground truth, not character by character, and the methodology and document mix are published in full on the accuracy page. On your own documents the number is established during the Proof of Value and fixed in writing before it begins.

What is covered

Seven themes, in the order a security review reaches them

If you are assembling an internal assessment, work down the page. The data and privacy section answers the questions that decide whether the rest of the evaluation happens at all.

Data and privacy

Training, controller and processor roles, retention, phone-home behaviour and what we can see during support.

Deployment and operations

Air-gapped operation, deployment models, hardware and sizing, offline updates, high availability and support.

Accuracy and evaluation

What the 99% figure means, hard documents, Arabic handling, low-confidence behaviour and how to run a fair bake-off.

Commercial

The fixed annual licence, reprocessing cost, how the paid Proof of Value works, renewals and who you contract with.

Security and compliance

CBUAE, UAE PDPL, the Health ICT Law, GDPR scope, what we will not claim, and the security pack under NDA.

Models and model licensing

The model manifest, who holds the weights, licence terms for offline commercial use, tuning and how updates arrive.

Integration and engineering

Document intake, output formats, retrieval pipelines, human review workflows and supported file types.

Data and privacy

The questions a security team asks first, because the answers decide whether the rest of the evaluation is worth running.

No. Your documents, extracted fields and any corrections your reviewers make never leave your deployment and are never used to train, fine-tune or improve any model, for you or for anyone else. In an on-premise install there is no mechanism for it: no collection pipeline, no outbound call, nothing to disable. It is also excluded contractually, so you have both the architecture and the clause.

You are the controller. Because DocxIntel runs inside your own infrastructure, we are not a processor for the documents you feed it — we never receive them. That normally removes the processing agreement, the sub-processor list and the cross-border transfer assessment from scope, and reduces the DPIA to an internal-only exercise.

Entirely according to your own storage policy. DocxIntel reads from the storage you already control and writes back to it; it holds no vendor-side copy, no cache in our infrastructure and no archive we could be compelled to produce. Retention, deletion and legal hold stay under the controls you already operate.

No. There is no licence check-in, no usage beacon and no crash reporting that requires an outbound connection, because the product is designed to run on networks that have none. You can verify this the way we would: deploy into an isolated segment, remove the default route, process a batch, and read your own egress and DNS logs.

Only if you decide to show us. Diagnostics are produced as log bundles your team reviews and releases, so you control exactly what leaves. Where deeper investigation is needed, sessions are escorted and scheduled by you, and for air-gapped environments on-site attendance is available instead.

Deployment and operations

What the infrastructure team needs to know before agreeing to host it: air-gap, hardware, updates, availability and support.

Yes. Air-gapped is a first-class deployment mode rather than a hardened variant of a cloud product. The models ship inside the bundle and load from local storage, and there is no fallback call to an external service when a page is difficult. A disconnected install performs identically to a connected one.

Four: air-gapped inside an isolated network, private cloud inside your own tenancy, managed single-tenant where we operate a dedicated environment for you, and an evaluation sandbox for assessment. All four run the same platform with the same feature set — the difference is where the boundary sits and who operates it.

Most production deployments run on a GPU-equipped server, and throughput scales horizontally by adding workers. A CPU-only profile exists for lower-volume, edge and evaluation deployments. It runs on bare metal, virtual machines or Kubernetes. Sizing is done against your real document mix and volume during the Proof of Value rather than from a generic reference specification.

Updates ship as signed, versioned offline packages that you import manually — nothing pulls from the internet on its own. Each release carries a changelog and a re-validation report, you choose when to promote it, you can validate in a sandbox first, and you can roll back to the previous model version.

Workers are stateless and horizontally scalable, the processing queue is durable, and the reference architectures cover multi-node and multi-site layouts. Because everything runs in your environment, recovery objectives are set by your own backup and restore standards, and we size the design against them during deployment rather than after an incident.

Accuracy and evaluation

What our published number means, what it does not promise, and how to measure the only figure that matters — the one on your documents.

It is 99% field-level accuracy on our published benchmark set, measured field by field against a human-adjudicated ground truth — not character-level accuracy, which flatters every vendor. The document mix, the adjudication method and the per-field breakdown are published alongside the number, because an accuracy figure without a methodology is a marketing figure.

Nobody can honestly promise that from a benchmark, ours included. What we do instead is measure it on your documents during the Proof of Value and agree the threshold in writing before we start. If the deployment misses the agreed threshold, you keep the report and owe nothing further.

Phone photographs, handwritten Arabic, stamps over printed text, multi-generation faxes and multi-page packets that need auto-splitting are treated as ordinary inputs. Deskew, dewarp, perspective correction, glare suppression and contrast normalisation run before recognition, and stamp ink is separated from the printed text beneath it, so a seal across an invoice total does not corrupt the total.

As a primary script, not a language pack. Right-to-left reading order, Arabic-Indic and Western numerals, diacritics, handwritten Arabic, and pages mixing Arabic body text with English headers or Latin product codes are resolved at the layout and reading-order layer. Reading order is established per block, so a bilingual table does not collapse into one scrambled string.

It is flagged rather than silently guessed. Every extracted value carries a page number, bounding-box coordinates and its own confidence score — not a document-level average — so your review queue receives the small proportion that genuinely need a human, and reviewers land on the exact region of the page rather than the whole document.

Assemble 300 to 500 real documents in the mix you actually receive, including the failures. Adjudicate the ground truth once, by hand. Run every candidate against the same set and score each field independently. Then price each candidate at your true annual volume, including reprocessing. Demos on vendor-supplied documents rank everyone as excellent and tell you nothing.

Commercial and licensing

How the fixed annual licence works, what the paid Proof of Value covers, and what happens at renewal.

As a fixed annual licence sized by deployment footprint — the environments, capacity and capabilities you run — never by page. Volume can double without the licence moving. This is the deliberate inverse of metered document AI, where the cost of reading a page never falls and the bill grows with your archive.

Nothing extra. The only constraint is the GPU hours on hardware you already own. This is the practical difference the licence model makes: under per-page pricing, reprocessing several million archived pages after a model improvement is a budget request, so most organisations simply never do it.

It is a paid engagement of 30 to 45 days, at a fixed fee, deployed inside your own environment on your own documents. The accuracy threshold that counts as success and the price you would pay on conversion are both agreed in writing before it begins. If the threshold is missed, you keep the full report and owe nothing further.

Because a real deployment inside a regulated environment involves our engineers, your infrastructure team and your security function for several weeks. A paid pilot gets both sides to commit properly and scope honestly. It also lets us fix the accuracy threshold contractually, which a free proof of concept never does.

Renewal covers support, updates and the licence to continue operating the deployment. Because the software and the model artifacts are already inside your data centre, the end-of-contract position is a licensing question rather than a service cut-off — and it is one you should settle in writing before signature, not after. Ask for the continuity clause explicitly.

With BizfyLabs, the engineering company based in Dubai that builds DocxIntel. DocxIntel is a product line rather than a separate legal entity, so the pilot agreement, the licence and the support relationship all sit with one counterparty.

Security and regulatory compliance

CBUAE, UAE PDPL, the Health ICT Law and GDPR — plus a plain answer on the certifications we do not hold.

On-premise and air-gapped deployment addresses the data-residency and outsourcing expectations directly, because customer documents never leave the licensed environment and there is no third-party processor in the path. The assessment shifts from evaluating an external service provider to evaluating software running inside your own controls.

Both place localisation and control obligations on exactly the categories of document that most need automating. Because processing happens inside your infrastructure, there is no cross-border transfer to assess and no external processor to notify. You remain the controller throughout, which is the position both regimes are written around.

It removes the hardest parts of the analysis rather than answering it for you. With no transmission to a third party there is no international transfer mechanism to justify and no sub-processor chain to disclose, and data-subject rights are exercised against systems you already control. Your obligations as controller are unchanged; the scope of the assessment is smaller.

We do not claim certifications we do not hold, and you should treat any vendor differently if they are vague on this question. What we can describe is how the architecture supports your own certification: the controls are designed against recognised frameworks, and because the deployment sits inside your environment it inherits the controls you are already audited on.

Yes, under NDA. It covers the architecture, the data-flow diagram, the control mapping, the model manifest and the update and support model. The data-flow diagram is the most useful part: during a pilot you can test it against your own firewall, DNS and proxy logs rather than taking it on trust.

Every extracted value keeps its page number, bounding-box coordinates and confidence score, and element identifiers are stable across re-runs, so a reviewer or an inspector can trace any figure back to the pixels it came from and diff behaviour between model versions. Audit trails live in your environment, under your retention policy.

Models and model licensing

Which models ship in the bundle, who holds the weights, what the licences permit and how new versions arrive offline.

Every deployment ships with a fixed, disclosed set of models covering OCR, layout, classification, extraction and the language model behind Ask. You receive a full manifest on day one naming each model, its version, its licence and its purpose — no black boxes and no silent upgrades.

Yes. The open-weight models arrive inside the deployment bundle as artifacts on storage you control, rather than being pulled from a vendor registry at runtime. That is what makes air-gapped operation possible at all, and it is what keeps the deployment functioning independently of the network and of the vendor relationship.

Yes, and you get the licence text rather than just the licence name so your legal team can verify it against your own position. There are no royalty or per-inference model fees layered on top of the DocxIntel licence.

Tuning is done as part of deployment and during the Proof of Value, against your document mix and your field definitions, entirely inside your environment. Nothing derived from your documents is transmitted to us or incorporated into any model we ship to anyone else.

It arrives in a signed offline update package with a changelog and a re-validation report, and you decide when to promote it. Because there is no per-page charge, you can then re-run historical documents through the improved model as a scheduling decision rather than a budget one — and the previous version stays available for rollback.

Integration and engineering

How documents get in, what comes out, and how DocxIntel fits alongside the systems and pipelines you already run.

From storage you already control: watched folders, S3-compatible object storage, your document management system, an intake mailbox, or a direct API call. Nothing is uploaded anywhere — the service reads from where your documents already live and writes results back to where your pipeline expects them.

Structured JSON containing blocks, tables, tokens, coordinates and confidence scores; reading-order markdown with table structure preserved for retrieval pipelines; annotated PDFs for review and sign-off; and per-table CSV or XLSX exports with merged cells resolved. Everything is returned over your own API inside your network.

Yes. The reading-order markdown and the coordinate-bearing JSON are the shapes retrieval frameworks expect, so most teams keep their orchestration, chunking and embedding layers and change only the parsing call. Citations can point back to the exact page region, which is what makes an answer reviewable.

Low-confidence fields are flagged with the page number and coordinates that produced them, so exceptions can be routed into the case-management or workflow tooling your reviewers already use rather than a separate vendor portal. Teams that want a supplied review interface out of the box should weigh that honestly when comparing platforms.

Native, scanned and hybrid PDFs; JPEG, PNG, TIFF, BMP, WEBP and HEIC images; DOCX, XLSX and PPTX plus the legacy Office formats; EML and MSG email with attachments; ZIP archives expanded in place; and Group 3/4 fax TIFFs including multi-generation copies. Coverage for your specific mix is confirmed in writing during the Proof of Value.

Keep reading

Deployment models→Air-gapped, private cloud, managed single-tenant and evaluation sandbox, in detail.Security architecture→The data-flow diagram, control mapping and how the design supports your own certification.Accuracy benchmark→Document mix, methodology and per-field results behind the 99% field-level figure.Pricing→Why the licence is fixed and sized by deployment footprint rather than metered per page.Model licences→The named models in the bundle and the licence terms attached to each of them.Compare approaches→Metered cloud API, BYOC on an enterprise tier and true on-premise, compared fairly.

If your question is not here, ask it directly.

We would rather answer a hard question in the first conversation than discover it in a security review three months later. Bring the objection you think will kill the project.

Talk to an engineer
How the Proof of Value works
  • No per-page metering
  • Runs in your environment
  • Written accuracy threshold
DocxIntel Logo

A product of BizfyLabs

Document intelligence that never leaves your building. Analyse, identify, classify, map, modify and ask — inside your own infrastructure.

BizfyLabs on LinkedInDocxIntel documentationBizfyLabs

Product

  • Capabilities
    • Analyse
    • Identify
    • Classify
    • Map
    • Modify
    • Ask
  • Accuracy benchmark
  • Pricing
  • Proof of Value

Technical

  • Deployment models
  • Reference architectures
  • Sizing & throughput
  • What's in the box
  • Security posture
  • Model licences
  • Documentation
  • API reference

Solutions

  • All industries
  • Insurance & TPAs
  • Healthcare
  • Banking & finance
  • Government
  • Legal
  • Energy & logistics

Compare

  • Compare approaches
  • LlamaParse alternative
  • Docsumo alternative
  • On-premise document AI

Company

  • About DocxIntel
  • FAQ
  • Partners
  • BizfyLabs
  • Careers
  • Contact

© 2026 BizfyLabs FZC LLC. All rights reserved.

DocxIntel™ is a product of BizfyLabs FZC LLC.

  • Privacy Policy·
  • Terms of Service·
  • Data Processing Addendum·
  • Acceptable Use·
  • Model Licences·
  • Security·
  • Cookies

Registered in the United Arab Emirates. Delivery partner: Bizfy Solutions LLP, Indore, India.

DocxIntel