Skip to content
All products

DocIntel · Document extraction

Turn any document into structure, in your languages, behind your firewall.

DocIntel detects layout, runs OCR across many languages including right-to-left and Devanagari scripts, and extracts tables and figures to clean JSON, XML, or a web viewer. It runs entirely on-premise.

The problem

Regulated organisations have warehouses of multilingual documents they can't send to the cloud, and generic OCR loses the tables, the structure, and the languages that matter.

How it works

From input to a result you can stand behind.

  1. 01

    Detect layout

    Identify titles, text, tables, and figures as regions, each with a confidence score.

  2. 02

    OCR in-language

    Recognise text across many languages, right-to-left and Devanagari scripts included.

  3. 03

    Extract & caption

    Pull tables into structured form and describe figures in natural language.

  4. 04

    Export anywhere

    Hand off clean JSON, XML, or an interactive web viewer, on your terms.

What's inside

Built for the work it does every day.

Layout detection

Title, text, table, and figure regions, each with a confidence score.

Multilingual OCR

Recognition across many languages, right-to-left and Devanagari scripts included.

Table extraction

Structured tables exported to Markdown and HTML.

Figure captioning

Natural-language descriptions of charts, diagrams, and images.

Three quality modes

Fast, Balanced, and Accurate: trade speed against fidelity per job.

On-premise deployment

Runs entirely inside your network, on a single GPU node.

A closer look

On-premise · air-gapped · single-node · multilingual
Extraction mode
AdvancedMin confidence 60%·Include XML

At a glance

Outputs
JSON, XML, web viewer
Languages
Multilingual, RTL + Devanagari
Deployment
Docker · single NVIDIA H100
Recognition
AI Grand Challenge 2025 finalist

Work with us

Turn any document into structure, in your languages, behind your firewall.