H1ST-AI-DBD-001 · Task 1.1
Your whole study waits on one manual read
When the protocol read is slow and inconsistent, every deliverable after it inherits the delay and the ambiguity. The Protocol Intelligence Agent reads a protocol in any format — PDF, Word, or scanned image via OCR — and turns it into structured, machine-readable elements, each linked back to its source, so the rest of your build starts from one agreed foundation instead of a fresh re-read.
The grind you know
A new protocol lands and the whole team stops to read it — again. You're hunting through hundreds of pages for objectives, inclusion and exclusion criteria, procedures, visit schedules, and endpoints, keying them out by hand, knowing a colleague reading the same document will pull something slightly different. Nothing downstream can start until this is done, and it always takes longer than anyone planned.
Part of the Database & Forms Design family
Protocol to CRF, ICF, SDTM & coding.
When the protocol read is slow and inconsistent, every deliverable after it inherits the delay and the ambiguity. The Protocol Intelligence Agent reads a protocol in any format — PDF, Word, or scanned image via OCR — and turns it into structured, machine-readable elements, each linked back to its source, so the rest of your build starts from one agreed foundation instead of a fresh re-read.
Drop in a PDF, Word file, or scanned document — OCR handles the rest.
Objectives, eligibility, procedures, visits, and endpoints are identified and organized.
Confirm the source-linked extraction; it feeds CRF, consent, and SDTM design automatically.
Capabilities
Parses PDF and Word natively and handles scanned protocols through Paddle OCR, supporting documents up to 500 pages.
Identifies study objectives, inclusion/exclusion criteria, procedures, visit schedules, efficacy endpoints, safety assessments, and laboratory parameters.
Reconstructs the visit-by-visit schedule of assessments as structured data that feeds CRF design and validation.
Every extracted element links back to its location in the source protocol, so reviewers can verify rather than re-read.
In practice
Where the agent shows up in the day-to-day of a live trial — the moments the grind usually lives in.
Your partner sends the protocol as a flat scanned PDF with no digital text underneath. Instead of retyping it, you drop it in and OCR reads all 300 pages, turning objectives, eligibility, and the visit schedule into structured elements you can actually work from.
You have watched two experienced colleagues read the same protocol and pull slightly different eligibility criteria. Here every extracted element links back to its exact location in the source, so the team verifies against one grounded foundation instead of re-reading and reconciling by hand.
You are about to start CRF design and need the visit-by-visit schedule of assessments as clean, structured data rather than a table buried on page 40. The agent reconstructs it automatically and hands it straight to CRF Designer so the build starts from an agreed structure.
Proof
Figures shown are pre-launch targets based on internal benchmarks, not guaranteed outcomes.
Paddle OCR
OCR
Gemini / Anthropic LLM
LLM
Who it's for
Turns a protocol into CDASH-compliant eCRFs with edit checks and validation rules.
Learn moreGenerates ICH-GCP-compliant informed consent forms at an 8th-grade reading level.
Learn moreAuto-generates SDTM annotations and the annotated CRF, Pinnacle 21-clean.
Learn morePeace of mind
Every output is generated inside a validated, audit-ready platform, kept under human-in-the-loop control, and mapped to the regulatory and CDISC standards this agent supports.
Walk through it on your own workflow with a clinical-trials expert — no pressure, no obligation, and honest answers, including on the limits.