Comment from Humane Intelligence

AnonymousSupportAdvocacy
Summary: Humane Intelligence, a 501(c)(3) nonprofit organization, supports the FDA's pilot program for AI-enabled optimization of early-phase clinical trials as a necessary step toward modernizing clinical systems. They argue that the FDA should move beyond general AI risk frameworks to establish prescriptive, domain-specific evaluation criteria that ensure data integrity, algorithmic rigor, and equitable representation across diverse patient populations.
Humane Intelligence — Comment Summary Docket No. FDA–2026–N–4390: AI-Enabled Optimization of Early-Phase Clinical Trials Pilot Program Humane Intelligence (HI) submits this comment to provide a public health and sociotechnical evaluation perspective on the FDA's proposed pilot program. As an organization dedicated to context-aware evaluations of algorithmic systems, our perspective centers directly on public health imperatives and the downstream, system-wide impacts of automated technology deployment in clinical development. Early-phase trials represent a critical, high-stakes bottleneck in drug development, establishing the fundamental safety profiles, biomarker parameters, and stratification logic that govern all subsequent clinical phases. While HI strongly supports this initiative as a necessary step toward modernizing legacy clinical systems, we urge the Agency to ensure that the pilot phase operates as a lean, rigorous learning environment by introducing pragmatic, up-front evaluation criteria. Our full attached comment outlines four primary structural recommendations: Translation of General-Purpose Meta-Frameworks: The pilot’s reliance on passive alignment with the general-purpose, domain-agnostic NIST AI Risk Management Framework (AI RMF) is operationally insufficient to protect clinical data integrity. By design, the NIST framework is a non-prescriptive blueprint meant to guide internal organizational cultures; it does not provide ready-to-use technical standards or operational thresholds. Deferring to it passes the regulatory burden down to individual commercial vendors, guaranteeing a highly fragmented technical execution across the pilot. The FDA should use this pilot to actively translate abstract NIST taxonomies into a prescriptive, domain-specific framework tailored explicitly for clinical trials, integrating stress-test requirements directly into the initial application process. Algorithmic Rigor and Representative Integrity in Small Cohorts: Layering predictive matching algorithms, biomarker selection software, or virtual control constructs onto early-phase trials introduces an exceptionally high risk of overfitting. If these tools overfit to the clean, uniform historical records of elite academic research centers, they will fail to generalize. This dynamic risks hard-coding historical public health exclusions into the drug development pipeline. The FDA must look past overall model accuracy and mandate subgroup-specific sensitivity metrics to evaluate system generalizability where real-world patient data is thinnest. Technical Uniformity and Data Provenance Across Stakeholders: AI architecture is transitioning toward dynamic, multi-model workflows. This introduces severe vulnerabilities regarding data provenance, model drift, and data residency when sensitive clinical text is routed across infrastructure networks. Leaving reporting metrics to the discretion of individual investigator sites will mask localized operational failures at less-resourced regional satellite sites. The FDA should require uniform evaluation rules and standardized reporting modules that bind technology vendors, biopharma sponsors, and clinical trial sites to a singular, verifiable data provenance standard. Operational Definition of Real-Time Review: Real-time data review must be operationally defined on both sides of the regulatory platform. If review windows are structured strictly around pure data velocity, the ingestion pipeline will favor elite health IT architectures and structurally bias the regulatory view. Furthermore, because early-phase clinical data evolves rapidly during dose escalation, the FDA must establish internal operational guardrails, including strict review frequencies and a review continuity index, to ensure that automated data speed does not compromise scientific oversight or cause reviewer fragmentation. Please see our attached document for detailed responses to Questions A.1.c, A.2.a, A.3.a, A.4.b, A.4.c, A.5.c, B.3.a, B.4.a, B.5.a, and B.5.e.

View on Regulations.gov