Security & data protection
AES-256 encryption, role-based access, PHI de-identification, enterprise-grade Google Cloud hosting, and 25-year eTMF archiving — the security architecture that lets an AI platform touch source data, consent, and the Trial Master File safely.
Sound familiar?
It’s six weeks to your sponsor’s audit, and a 300-row security questionnaire just landed on your desk — and the auditor wants evidence, not adjectives.
And the part no one says out loud:
It doesn't have to be this way
Security here is architectural, not aspirational. Every control is evidenced, every access is logged, and the answers to the hard questions are already written down — so you walk into the review ready instead of scrambling.
Security by architecture
Health1st AI handles some of the most sensitive data in medicine — patient information, source documents, and inspection-critical records. Security is not a feature bolted on top; it is the foundation every agent runs on, and it is the same foundation our compliance certifications are audited against.
Below are the five pillars of that foundation. Each one is continuously monitored and evidenced as part of our SOC 2 and ISO 27001-certified information-security management system.
Encryption
Trial data — source documents, PHI, datasets, and the eTMF — is encrypted everywhere it lives and everywhere it moves.
All stored data is encrypted with AES-256, including databases, object storage, backups, and the document repository. Encryption keys are managed in a dedicated key-management service with rotation and strict access controls.
Every connection — user sessions, API calls, and integration sync with your EDC, randomization, and document systems — is protected with TLS, so data is never exposed on the wire.
Keys are separated from the data they protect, access to them is logged, and rotation is enforced, so a single credential can never unlock the whole platform.
Access control
Every user, agent, and integration sees only what their role requires — and every access is authenticated and logged.
Permissions map to clinical roles — data manager, biostatistician, monitor, medical writer, QA, project manager — so people reach the studies, documents, and functions they own and nothing more.
Access requires strong authentication with support for single sign-on (SSO) and multi-factor authentication, and sessions are governed by timeout and re-authentication policies.
Study- and site-level segregation keeps data partitioned, while a 21 CFR Part 11 audit trail records who accessed or changed what, when, and why.
Privacy
The platform is designed to process the minimum identifying data necessary, and to strip identifiers wherever downstream work does not need them.
Protected health information is de-identified before it flows into AI processing, analytics, or any integration destination that does not require identifiers — reducing exposure without slowing the workflow.
Agents operate on the least data required for the task, and PHI is compartmentalized so it is not duplicated across systems that have no need for it.
De-identification, role-based access, encryption, and lawful-basis controls support HIPAA and GDPR, with Business Associate Agreements and EU data-residency options available.
Hosting
The platform runs on Google Cloud Platform, inheriting enterprise-grade physical, network, and infrastructure security.
GCP provides certified data centers, network isolation, DDoS protection, and infrastructure controls audited to global standards, underpinning the platform's own SOC 2 and ISO 27001 posture.
Regional deployment options support data-residency requirements for global trials, including EU residency for GDPR-governed studies.
Redundancy, automated backups, and disaster-recovery processes protect availability so study operations continue through disruptions.
Retention
Regulations require Trial Master File records to be retained for decades. The platform keeps the eTMF complete, legible, and retrievable for the full retention period.
The electronic Trial Master File and its records are retained for up to 25 years, meeting ICH E6(R3) and regional retention requirements without migrating data off to fragile offline media.
Archived records stay Enduring and Available under ALCOA+ — versioned, access-controlled, and retrievable for inspection with their audit trail intact.
Fixity checks, immutable audit trails, and controlled access preserve the integrity and attributability of records across the entire retention lifecycle.
Where your data lives
When an auditor asks where your source documents, PHI, datasets, and eTMF actually sit, you should have a single, provable answer. Here is how the platform is partitioned, where data resides, and how it moves.
Your program runs in a logically isolated tenant, and studies and sites are segregated within it. One sponsor's data is never commingled with another's, and access is scoped tenant-first at every layer — storage, compute, and the document repository.
Choose where your data is stored. Regional Google Cloud deployment supports data-residency requirements for global trials, including EU residency for GDPR-governed studies, so cross-border transfer rules are handled by design rather than by exception.
Every category of data — source values, randomization events, documents, SDTM/ADaM datasets, and the eTMF — has a defined home and a traceable path. As data moves between the platform and your systems of record, each hop is authenticated, permissioned, and recorded, so lineage is reconstructable end to end.
Data is retained on the eTMF's regulatory schedule and your contract, then disposed of through controlled, logged processes. You keep ownership throughout, and on exit your records are exported in full — nothing is held hostage.
AI safeguards
An AI platform working on source data and the Trial Master File has to clear a higher bar than a generic tool. These are the controls that make agentic AI trustworthy in a regulated trial.
Your trial data, documents, and PHI are never used to train foundation models or shared across customers. Models are used for inference only, under contractual and technical controls that keep your data yours.
Agents generate against your actual study artifacts using retrieval — not open-ended recall. Outputs are grounded in source and citable back to the document or value they came from, so a reviewer can trace every claim.
PHI is de-identified before it reaches AI processing wherever identifiers are not required, minimizing exposure without breaking the workflow.
Agents propose; qualified people dispose. Every regulated output routes through human review and 21 CFR Part 11 electronic signature before it becomes part of the record — the model never has the last word.
AI use is managed under an ISO 42001 AI-management-system framework, with documented model governance, monitoring, and controls sitting on top of the platform's ISO 27001 and SOC 2 posture.
Agents reach only the data their task and the operator's role permit. The same RBAC and audit trail that govern people govern the agents acting on their behalf, so nothing runs unsupervised or unlogged.
Incident response & continuity
Trust is not only about preventing incidents — it is about how quickly and honestly they are handled, and whether your inspection-critical records stay available through disruption.
The environment is monitored continuously for anomalous access and threats, with alerting tied into the security operations that back our SOC 2 program.
A documented incident-response plan governs detection, containment, investigation, and notification. If an event affecting your data occurs, you are informed on contractual and regulatory timelines — not left in the dark.
Automated backups and tested disaster-recovery processes on Google Cloud protect against data loss, with recovery objectives that keep study operations moving through disruption.
Redundant infrastructure and documented continuity procedures keep the platform — and access to your inspection-critical records — available when you need them, including during an active inspection.
How it fits together
The controls on this page are the technical half of a broader trust model. See how they map to the standards regulators expect and the systems you already run.
We'll share our architecture, controls, and certification status under NDA and answer your security and privacy questionnaires in detail.