AI document engine · domain-agnostic

Your documents become searchable knowledge

datamida.ai is an AI document engine: it ingests documents from a web interface or via REST API, classifies them, automatically extracts metadata and makes them queryable in natural language. One engine, any sector, with the data perimeter always under your control.

The problem

Why does your document estate stay unused?

Contracts, financials, ID documents and case files are archived in folders and back-office systems, but they are not queryable: the knowledge exists, it is just unreadable at scale. Every question becomes a manual search and every extraction work redone. Without classification and metadata, an archive grows in volume but not in value.

  • Documents siloed, never pooled together
  • Metadata missing or filled in by hand
  • No natural-language answers on the content

What it is

What is datamida.ai?

datamida.ai is an AI document engine that turns archives of unstructured documents into a queryable knowledge base. It ingests files from a web interface or via REST API, automatically recognizes the type of each document, extracts metadata as JSON and indexes it for natural-language search.

It is not a standalone app to adopt separately: it is an engine your existing back-office software integrates via API, so classification, extraction and search live inside your current processes. Each client gets a dedicated single-tenant instance and chooses, through configuration, where data processing happens. The product is developed by Pixel Service & Consulting Srl.

How it works

How does the document processing pipeline work?

Every document, regardless of how it enters the system, runs the same three-stage pipeline: parse → classify → extract. During parsing the content is read and normalized from PDFs, images and scans. During classification the AI recognizes the document type — for example a financial statement, a contract or an ID document. During extraction it produces structured fields as JSON, ready to be indexed. Once complete, the document is indexed and instantly queryable.

  1. parse
    Ingest
    PDFs, images, scans
  2. classify
    Classify
    Document type recognized
  3. extract
    Extract metadata
    Structured fields as JSON
  4. indexed
    Indexed
    Searchable instantly

Three ways to feed the same engine

The same pipeline powers three ingestion modes, so you can start with manual use and move to full automation without changing engine.

01 · UI

Interactive ingestion with operator

The AI proposes the extracted metadata, the operator confirms or corrects it, then the document is indexed. It is the ideal path for maximum data quality when human supervision is required. analyze → confirm

02 · API

Headless extraction via API

Synchronous extraction with no persistence: you send the document, metadata returns as JSON and the downstream system decides what to do. It serves machine-to-machine flows that need no storage on datamida.ai's side. extract

03 · AUTO

Automatic ingestion with validation

Three validation gates — document type, required fields and confidence level — replace the operator and decide whether to index or flag. It enables full automation over large volumes. ingest

Domain-agnostic

What sectors does a domain-agnostic engine work in?

datamida.ai is domain-agnostic: no industry-specific logic is written into the code. The engine adapts to the client's domain through simple configuration, so the same system serves very different contexts without rewriting the software.

You change the configuration — document types, fields to extract, validation rules — not the engine. This lowers the cost of adoption in a new domain and keeps a single codebase to maintain and update.

config not code
  • Finance & guarantees
    financials, filings, cases
  • Legal
    deeds, opinions, contracts
  • Healthcare
    reports, records, IDs
  • Public admin
    protocols, correspondence

Deployment & GDPR

How do you choose the level of GDPR compliance?

Compliance is a configuration choice, not a constraint imposed by the product. Each client has its own dedicated single-tenant instance and decides where data processing happens by picking one of three deployment modes.

● live

Full On-Premise

AI processing runs entirely on-premise, with no calls to external clouds. Ideal for strict security policies and data-residency requirements: zero data leaves the corporate perimeter.

◐ coming soon

Hybrid EU

AI processing within the European Union via enterprise cloud providers, with SLAs and scalability, keeping data in the EU. It balances European data residency and managed performance.

● live

Direct API · DPA

Frontier AI providers with DPA and standard contractual clauses. The fastest route to a proof-of-concept, at top model quality and minimal time-to-deploy.

Integration

How does datamida.ai integrate with existing systems?

datamida.ai integrates via REST API and lives inside the back-office software the client already uses, not beside it. Beyond interactive use with an operator, it exposes synchronous headless extraction and automatic ingestion for machine-to-machine flows.

The instance is single-tenant and isolated, answers stream in real time, and system-to-system integrations are ready to use, so document functions become part of existing processes.

  • Native integration
  • API-first
  • Real-time
  • Single-tenant
  • Secure by design
  • Scalable

FAQ

Frequently asked questions about datamida.ai

What is datamida.ai?

datamida.ai is an AI document engine: it ingests documents from a web interface or via REST API, classifies them automatically, extracts metadata as JSON and makes them searchable in natural language. It is not a standalone app but an engine your back-office software integrates via API. Each client gets a dedicated single-tenant instance, and it is developed by Pixel Service & Consulting Srl.

In which sectors can I use datamida.ai?

datamida.ai is domain-agnostic: no industry logic is written into the code. The engine adapts to your domain through configuration, so the same system serves different contexts such as finance and guarantees, legal, healthcare and public administration without rewriting the software.

How does document search work?

Search runs across three complementary channels from the same catalog: semantic RAG with clickable citations to the sources, hybrid search combining metadata filters and semantic similarity, and direct SQL over the structured catalog for precise extractions and reports.

Which documents and formats does datamida.ai accept?

datamida.ai ingests documents in PDF, image and scan formats. Content is read and normalized during the parsing stage; the document type is then recognized during classification, regardless of the structure of the source file.

Is datamida.ai GDPR compliant?

Yes. Compliance is a configuration choice: three deployment modes are available — Full On-Premise with entirely local AI processing, Hybrid EU with data residency in the European Union, and Direct API with DPA and standard contractual clauses. Each client has its own dedicated single-tenant instance and sets the data perimeter through configuration.

How do I integrate datamida.ai with my back-office system?

datamida.ai integrates via REST API. Beyond interactive use with an operator, it offers synchronous headless extraction (extract) and automatic ingestion with validation gates (ingest) for machine-to-machine flows, with real-time answer streaming and system-to-system integrations.

Is datamida.ai a standalone product or an engine to integrate?

datamida.ai is a document engine, not a standalone app. It lives inside the software the client already uses, integrating via REST API on a dedicated single-tenant instance. Classification, extraction and search become part of existing processes instead of requiring a separate tool.

Let's turn your documents into a knowledge engine

Contact us