Skip to content
Hyperfluid 2.0 is live console.hyperfluid.cloud
Document Intelligence

Document Agents

Your documents become queryable data

Pipelines that turn your PDFs, emails and attachments into structured, queryable data: OCR, classification, context enrichment and routing. OCR and classification run on a model served in your own cluster, or on a European API if you prefer.

Document Agents Interface - Hyperfluid document processing pipelines

Specifications

OCR & Extraction
PDF, Word, images into structured data
AI Classification
Automatic labeling and sorting
Hippocampe
L0/L1 summaries, progressive search
Routing
To the right destinations
Pipelines
9 types available
Destinations
Iceberg tables, object storage

Use cases

Sort incoming documents

Invoices, quotes, letters: pipelines automatically classify and file every incoming document

Automatic sorting, folder by folder

Query your PDFs

OCR extraction turns your documents into structured data, queryable like any other table

From PDF to Iceberg table

Build a document memory

Hippocampe generates L0 and L1 summaries for progressive document search

Progressive search

In action

Classify the inbound document flow

A property management team receives invoices, quotes and letters by email every day

  1. Plug the email connector into the management inbox
  2. Enable a document classification pipeline
  3. Routing delivers each document to the right folder, property by property
  4. Track runs from the pipelines dashboard

No more manual sorting: every document lands classified in the right place

Make PDF archives queryable

A data team needs to make years of accumulated PDF archives usable

  1. Drop the archives into the organization's object storage
  2. Run an OCR extraction pipeline on the batch
  3. Structured data lands in Iceberg tables
  4. Query the archives in SQL from the console

Years of archives become queryable data

Key benefits

  • Your documents become queryable data
  • 9 ready-to-use pipeline types, tracked from a dedicated dashboard
  • Fed by your connectors: inbound email, Google Drive, object storage
  • Everything runs inside your sovereign perimeter

Ready to make your documents talk?

Discover how document agents turn your documents into actionable data.