H1ST-AI-DOC-006 · Task 6.2
Filing a document shouldn't be a judgment call
You shouldn't need to remember the whole Reference Model to file one document correctly. The agent inspects each document's content and layout to predict its TMF zone, section, and artifact under Reference Model v3.3.1, extracts key metadata such as site, country, dates, and version, and flags likely duplicates before they pollute the TMF, returning a confidence score for your review. Filing that once took minutes of manual judgment resolves in seconds against a consistent taxonomy.
The grind you know
Classifying TMF documents by hand is slow, and worse, it depends on which coordinator is doing it, so artifacts end up mis-filed and records orphaned. The metadata you key by hand carries the typos and gaps that later break your search and your completeness reporting. And the duplicates pile up quietly, inflating the inspection risk you can't yet see.
Part of the eTMF & Documents family
An eTMF that files and manages itself.
You shouldn't need to remember the whole Reference Model to file one document correctly. The agent inspects each document's content and layout to predict its TMF zone, section, and artifact under Reference Model v3.3.1, extracts key metadata such as site, country, dates, and version, and flags likely duplicates before they pollute the TMF, returning a confidence score for your review. Filing that once took minutes of manual judgment resolves in seconds against a consistent taxonomy.
The agent parses text and layout from PDFs, scans, and Office files to build a content representation.
A model maps the content to the Reference Model taxonomy and extracts metadata, scoring each prediction.
High-confidence results file automatically; uncertain or duplicate items are surfaced for reviewer confirmation.
Capabilities
Content and layout analysis assigns each document to the correct TMF zone, section, and artifact with a confidence score.
The agent pulls site, country, investigator, version, and key dates directly from the document to auto-populate index fields.
Fingerprinting and near-match comparison flag duplicate or superseded copies before they are filed.
Low-confidence predictions are routed to a human queue while high-confidence items proceed automatically.
In practice
Where the agent shows up in the day-to-day of a live trial — the moments the grind usually lives in.
A site returns a signed form as a scanned PDF. PaddleOCR converts it to text, and the agent predicts the correct TMF zone, section, and artifact under Reference Model v3.3.1, extracting site, country, and dates so it's indexed like a native file.
An ambiguous document scores below threshold. Rather than mis-file it, the agent routes it to a reviewer queue with its confidence score, while high-confidence items proceed automatically.
A coordinator uploads a copy of a document already on file. Fingerprinting and near-match comparison flag it as a duplicate or superseded copy before it inflates your completeness reporting.
Proof
Figures shown are pre-launch targets based on internal benchmarks, not guaranteed outcomes.
Paddle OCR
OCR
Gemini / Anthropic LLM
LLM
Veeva Vault
EDC
Who it's for
An intelligent dump-folder that auto-assigns metadata and files documents to the correct TMF location behind a review-and-approve gate.
Learn moreA BPMN-driven workflow engine that routes, approves, and lifecycles Trial Master File documents against the TMF Reference Model v3.3.1.
Learn moreDocument version control with change tracking, side-by-side comparison, and controlled archive management for the TMF.
Learn morePeace of mind
Every output is generated inside a validated, audit-ready platform, kept under human-in-the-loop control, and mapped to the regulatory and CDISC standards this agent supports.
Walk through it on your own workflow with a clinical-trials expert — no pressure, no obligation, and honest answers, including on the limits.