AI classification
Know what every file is before anything moves
Mergiva uses a large language model to place every discovered file in your taxonomy and score its sensitivity. You choose the model for each AI feature, the account it runs under and how personal data is handled.
- Claude Haiku 4.5 by default
- Your key or your gateway
- Personal data redacted first
- Model chosen per feature
Why classify first
The plan for a regulated estate starts with what is in it
In a life-sciences deal the first questions are not about bytes. Where are the clinical records? Which files hold personal or health data? What is competitively sensitive, and must stay behind the antitrust gate until Day 1? The answers decide how the estate is split into waves, who signs for each one, and what your privacy and quality teams need to see.
Mergiva answers them file by file. Once discovery has recorded and hashed every file, a large language model reads each file’s details, places it in your taxonomy and scores its sensitivity from 1 to 5. Every answer is stored with the model’s confidence and its reasoning, so a reviewer can see why a file was tagged the way it was.
You stay in charge of the AI. You choose the model and the account it runs under, how personal data is handled before a request leaves your cluster, and how much file content the model may see. If your policy is that no personal data leaves, set the guard to refuse any classification request that contains it.
How it works
How one file is classified
Six steps, and five of them run inside your cluster. The diagram shows the path a single file takes; the notes below say what happens at each step.
Step 1
Discovered file
Name, path and type
Step 2
Prompt from your taxonomy
Your categories and scale
Step 3
Personal data guard
Redact, block or warn
Step 4
Your chosen model
Claude, or your gateway
Step 5
Checked answer
Only your tags accepted
Step 6
Stored with the file
Tags, score, confidence
Step 1
Discovered file
Name, path and type
Step 2
Prompt from your taxonomy
Your categories and scale
Step 3
Personal data guard
Redact, block or warn
Step 4The one outside call
Your chosen model
Claude, or your gateway
Step 5
Checked answer
Only your tags accepted
Step 6
Stored with the file
Tags, score, confidence
- 1
Discovered file
The discovery scan has already recorded each file’s name, folder path, content type and SHA-256. For files on a local or NAS source, the classifier also reads up to the first 4 KB.
- 2
Prompt from your taxonomy
The classifier builds the prompt from your taxonomy profile: every category with its definition, and the sensitivity scale. It asks for one answer in a fixed format.
- 3
Personal data guard
Before the request leaves your cluster, Microsoft Presidio looks for personal data. By default it redacts what it finds. You can make it refuse the call instead, or only log the finding.
- 4
Your chosen model
The request goes to the model you chose for classification, Claude Haiku 4.5 unless you set another. It travels to Anthropic under your own key, or through your own AI gateway.
- 5
Checked answer
The answer is parsed and checked. Tags outside your taxonomy are dropped, and the score must sit on your scale. If the answer fails the check or the call fails, the file is marked as a fallback with zero confidence, so no one mistakes it for the model’s view.
- 6
Stored with the file
The result is stored against the file: its category tags, a sensitivity score, the model’s confidence and reasoning, and whether it came from the model or the fallback.
What the model sees
Exactly what goes into the prompt
What the model receives depends on where the file lives. The prompt holds nothing else: no deal name, no hashes, no other files.
- Local folders and NAS shares
- File name, folder path, content type, and up to the first 4 KB of the file. Text files are sent as text; other files as a short hexadecimal excerpt.
- Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP
- File name, folder path and content type. No file content.
- Every request
- Your taxonomy’s categories with their definitions, the 1 to 5 sensitivity scale, and the format the answer must take. By default, detected personal data is redacted from the file’s details before the call.
You configure it
Your model, your account, your rules
Each setting below exists in the product today. The label under each one says where it is set.
Your model, feature by feature
Choose the model for each AI feature, classification included. Claude Haiku 4.5 is the default, and Claude Sonnet 4.6 and Claude Opus 4.6 are also offered. Behind your own gateway, you list the models it serves.
Where: AI Operations screen, by an administrator
Your own account, or your own gateway
Calls run under the API key you give the installation, so they fall under your own agreement with Anthropic. Or point Mergiva at your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. Every call goes only where you route it. A call your settings cannot serve is refused, never sent somewhere else.
Where: Installation settings
How personal data is handled
Redact detected personal data before the call (the default), refuse any request that contains it, or only log it. A second, optional check scans the model’s answer and redacts any personal data it repeats.
Where: Installation settings
Your taxonomy
The categories, their definitions and the names of the five sensitivity levels come from a taxonomy profile. The default is the pharma profile shown on this page. You can supply your own; the scale stays 1 to 5.
Where: A profile file, read when the service starts
How much content the model sees
Set how much of a local or NAS file goes into the prompt, from the default of 4 KB down to none.
Where: Installation settings
Rate limits and spend budgets
Cap requests and tokens per minute, concurrent calls and the size of a single request. Set daily or monthly spend budgets for each tenant. Both are checked before every classification call.
Where: By a platform or tenant administrator
A circuit breaker
When the model keeps failing, Mergiva stops calling it until it recovers, so an outage fails fast instead of stalling every scan. Calls refused for rate limiting are retried with a growing delay.
Where: Built in
Usage you can see
The AI Operations screen shows calls, tokens, cost, cache hits and latency for classification, with the model each call used.
Where: AI Operations screen
The default taxonomy
Eight categories and a five-level scale
These are the categories and definitions the model is given, exactly as the default pharma profile states them. Replace them with your own profile if your organisation classifies differently.
- GxP-Clinical
- clinical trial data, protocols, case report forms
- GxP-Manufacturing
- batch records, deviations, equipment validation
- GxP-Regulatory
- FDA/EMA submissions, 510(k), CMC, CTA, labeling
- PHI
- personal health information (named patients, MRNs, diagnoses)
- PII
- personal information without health context (HR records, contractor names)
- Financial
- revenue, accounting, internal financials
- Legal-Antitrust
- HSR, FTC/DOJ filings, antitrust counsel work product
- Other
- none of the above
Sensitivity, scored 1 to 5
- 1Public
- 2Internal
- 3Confidential
- 4Restricted
- 5Highly Restricted
A file can carry more than one category, such as clinical and health data together. It always carries exactly one sensitivity score.
Where results go
Classification shapes the rest of the deal
Scoping waves
The Data Estate Report
The Classification Explorer
Questions
What buyers ask about the AI
Which AI model does Mergiva use?
Claude Haiku 4.5 by default. Claude Sonnet 4.6 and Claude Opus 4.6 are also supported, and an administrator can choose the model for each AI feature. Behind your own gateway, you list the models it serves.
Can we use our own AI account?
Yes. The installation calls the model with the API key you give it, under your own agreement with Anthropic. Or route every call through your own AI gateway: you set its address, its key and any headers it needs, and it can speak the Anthropic or the OpenAI format. A call your settings cannot serve is refused, never sent somewhere else.
What exactly does the model see?
For files on local or NAS sources: the name, folder path, content type and up to the first 4 KB. For Amazon S3, Azure Blob Storage, Google Cloud Storage, MinIO, SharePoint Online and SFTP: the name, path and type only. By default, detected personal data is redacted before the call.
How can we check what the model decided?
Each result carries the model’s confidence and its reasoning, so a reviewer can see why a file was tagged. An answer that does not fit your taxonomy, or a failed call, is stored as a fallback with zero confidence, so it is never mistaken for the model’s view.
Does the model decide where files go?
No. The model tags files. People choose each wave’s destination, and two different people sign it. Transfers, signatures and evidence do not depend on the model.
Start with one deal.
Judge us on the ledger, not the demo.
1
Name the pair
Tell us the two systems you need to connect. We produce that pair’s evidence before the pilot starts.
2
Scan one estate
Run a Data Estate Scan in your own cluster. You get the PDF report and a classification your QA team can inspect.
3
Plan validation together
Evidence maps, the control inventory and test artefacts, executed with your QA team on your infrastructure.
4
Run the first wave
Two signatures, a verified transfer and a compliance report you can hand to an assessor.
Or write to contact@mergiva-ai.com.