Comment from Alexis Collier

AnonymousSupportAcademic
Summary: The commenter, a researcher specializing in AI reliability and clinical documentation, supports the pilot program but emphasizes the need to measure the underlying data environment's integrity alongside AI performance. They argue for disaggregated reporting by site and subgroup, pre-specified fairness analyses, and the inclusion of patient groups and investigators in governance to ensure safety and equity.
I appreciate the opportunity to comment on the proposed pilot program for AI-enabled optimization of early-phase clinical trials. My research focuses on the reliability, fairness, and governance of AI in real clinical settings, with particular attention to the documentation and data environment that AI systems depend on. My comments address Category A (Pilot Program Design) and Category B (Evaluation Metrics and Success Criteria), with emphasis on data integrity, generalizability, and trustworthiness. A.1 Scope and Focus. AI use cases for safety monitoring and participant screening should be prioritized, because both depend heavily on the quality of underlying clinical documentation and both carry direct safety and equity consequences. These are also the use cases where unreliable input data is most likely to cause harm, which makes them the most informative to study under controlled pilot conditions. A.3 Collaboration Models. Patient groups and investigators should have a defined role in AI governance from the start, not only at review. Early-phase trials in under-resourced and community settings are where AI tools are least validated and most consequential. Including the affected investigators and patient representatives in governance reduces the risk that efficiency gains accrue unevenly. B.3 Participant Safety and Data Integrity. FDA asks how to assess improvements in data completeness, accuracy, and consistency. I recommend that the pilot measure the reliability of the documentation environment itself, not only the outputs of the AI system. Documentation burden, charting delay, and missingness are measurable upstream signals that degrade data quality before any model is applied, and AI systems trained or operated on strained documentation inherit that weakness. Concretely, the pilot should capture baseline documentation-quality metrics (completeness, timeliness, and missingness) at each participating site before AI deployment, track whether AI-enabled approaches improve or worsen those metrics over time, and treat documentation-environment reliability as a covariate when interpreting AI performance, so that a model is not credited for performance that actually reflects a cleaner site. This makes data-integrity evaluation causal and site-aware rather than purely outcome-based. B.4 AI System Performance. FDA asks how performance should be evaluated across patient populations, trial sites, and therapeutic areas, and how model drift should be detected. I recommend that performance be reported disaggregated by site and by clinically relevant subgroups by default, never as a single pooled number, because aggregate performance can mask substantial degradation in specific settings. Drift monitoring should be continuous and tied to the documentation-environment metrics above, since changes in charting behavior are a common and overlooked driver of apparent model drift. B.5 Trustworthiness, aligned with the NIST AI Risk Management Framework. FDA asks what approaches should assess fairness across demographic and clinical subgroups, and what metrics apply to both sponsor-developed and proprietary systems. I recommend pre-specified subgroup fairness analysis with explicit, reported thresholds, evaluated before and during deployment rather than only at the end; evaluation of disparate performance and not only disparate outcomes, since a tool can appear fair on outcomes while performing worse for the subgroups it most affects; and, for proprietary systems, required reporting of standardized, model-agnostic performance and fairness metrics computed on held-out evaluation data, so that proprietary status does not prevent independent assessment of reliability and fairness. B.7 Qualitative Outcomes. Usability and workflow integration should be measured from the perspective of the clinicians and staff who carry the documentation and coordination load. If an AI tool speeds enrollment while increasing clinician burden, its net value is questionable and its adoption will not be durable. Stakeholder trust should be assessed with attention to whether the tool reduces or adds to the documentation work that already discourages clinician participation in research. Summary recommendation. The pilot will produce more credible and generalizable lessons if it measures the reliability and fairness of the documentation and data environment alongside the AI system, rather than evaluating model outputs in isolation. Documentation-environment metrics, site-disaggregated and subgroup-disaggregated reporting, and pre-specified fairness analysis are practical, near-term measures that support FDA's goal of faster early-phase development without sacrificing safety, data integrity, or equity. Thank you for considering these comments. Disclosure: I research documentation-derived measures of clinical workflow and am named inventor on U.S. provisional patent applications in this area. These comments reflect my research perspective.

View on Regulations.gov