Interactive ingestion with operator
The AI proposes the extracted metadata, the operator confirms or corrects it, then the document is indexed. It is the ideal path for maximum data quality when human supervision is required. analyze → confirm
AI document engine · domain-agnostic
datamida.ai is an AI document engine: it ingests documents from a web interface or via REST API, classifies them, automatically extracts metadata and makes them queryable in natural language. One engine, any sector, with the data perimeter always under your control.
The problem
Contracts, financials, ID documents and case files are archived in folders and back-office systems, but they are not queryable: the knowledge exists, it is just unreadable at scale. Every question becomes a manual search and every extraction work redone. Without classification and metadata, an archive grows in volume but not in value.
What it is
datamida.ai is an AI document engine that turns archives of unstructured documents into a queryable knowledge base. It ingests files from a web interface or via REST API, automatically recognizes the type of each document, extracts metadata as JSON and indexes it for natural-language search.
It is not a standalone app to adopt separately: it is an engine your existing back-office software integrates via API, so classification, extraction and search live inside your current processes. Each client gets a dedicated single-tenant instance and chooses, through configuration, where data processing happens. The product is developed by Pixel Service & Consulting Srl.
How it works
Every document, regardless of how it enters the system, runs the same three-stage pipeline: parse → classify → extract. During parsing the content is read and normalized from PDFs, images and scans. During classification the AI recognizes the document type — for example a financial statement, a contract or an ID document. During extraction it produces structured fields as JSON, ready to be indexed. Once complete, the document is indexed and instantly queryable.
The same pipeline powers three ingestion modes, so you can start with manual use and move to full automation without changing engine.
The AI proposes the extracted metadata, the operator confirms or corrects it, then the document is indexed. It is the ideal path for maximum data quality when human supervision is required. analyze → confirm
Synchronous extraction with no persistence: you send the document, metadata returns as JSON and the downstream system decides what to do. It serves machine-to-machine flows that need no storage on datamida.ai's side. extract
Three validation gates — document type, required fields and confidence level — replace the operator and decide whether to index or flag. It enables full automation over large volumes. ingest
Search
Search runs across three complementary channels that draw from the same catalog of indexed documents. Each channel answers a different type of question, from a natural-language request to precise extraction over structured fields.
Open questions on the content, answered with clickable citations pointing to the exact source.
Metadata filters and semantic similarity, combined automatically to balance precision and recall.
Direct querying of structured metadata for deterministic, repeatable extractions and reports.
Which financials have revenue above 100k?
I found 3 filings with revenue above €100,000. The most recent reports €142,000 in revenue.
Illustrative example of a RAG answer with source citations.
Domain-agnostic
datamida.ai is domain-agnostic: no industry-specific logic is written into the code. The engine adapts to the client's domain through simple configuration, so the same system serves very different contexts without rewriting the software.
You change the configuration — document types, fields to extract, validation rules — not the engine. This lowers the cost of adoption in a new domain and keeps a single codebase to maintain and update.
config not codeDeployment & GDPR
Compliance is a configuration choice, not a constraint imposed by the product. Each client has its own dedicated single-tenant instance and decides where data processing happens by picking one of three deployment modes.
AI processing runs entirely on-premise, with no calls to external clouds. Ideal for strict security policies and data-residency requirements: zero data leaves the corporate perimeter.
AI processing within the European Union via enterprise cloud providers, with SLAs and scalability, keeping data in the EU. It balances European data residency and managed performance.
Frontier AI providers with DPA and standard contractual clauses. The fastest route to a proof-of-concept, at top model quality and minimal time-to-deploy.
Integration
datamida.ai integrates via REST API and lives inside the back-office software the client already uses, not beside it. Beyond interactive use with an operator, it exposes synchronous headless extraction and automatic ingestion for machine-to-machine flows.
The instance is single-tenant and isolated, answers stream in real time, and system-to-system integrations are ready to use, so document functions become part of existing processes.
FAQ
datamida.ai is an AI document engine: it ingests documents from a web interface or via REST API, classifies them automatically, extracts metadata as JSON and makes them searchable in natural language. It is not a standalone app but an engine your back-office software integrates via API. Each client gets a dedicated single-tenant instance, and it is developed by Pixel Service & Consulting Srl.
datamida.ai is domain-agnostic: no industry logic is written into the code. The engine adapts to your domain through configuration, so the same system serves different contexts such as finance and guarantees, legal, healthcare and public administration without rewriting the software.
Search runs across three complementary channels from the same catalog: semantic RAG with clickable citations to the sources, hybrid search combining metadata filters and semantic similarity, and direct SQL over the structured catalog for precise extractions and reports.
datamida.ai ingests documents in PDF, image and scan formats. Content is read and normalized during the parsing stage; the document type is then recognized during classification, regardless of the structure of the source file.
Yes. Compliance is a configuration choice: three deployment modes are available — Full On-Premise with entirely local AI processing, Hybrid EU with data residency in the European Union, and Direct API with DPA and standard contractual clauses. Each client has its own dedicated single-tenant instance and sets the data perimeter through configuration.
datamida.ai integrates via REST API. Beyond interactive use with an operator, it offers synchronous headless extraction (extract) and automatic ingestion with validation gates (ingest) for machine-to-machine flows, with real-time answer streaming and system-to-system integrations.
datamida.ai is a document engine, not a standalone app. It lives inside the software the client already uses, integrating via REST API on a dedicated single-tenant instance. Classification, extraction and search become part of existing processes instead of requiring a separate tool.