AI Evaluation Systems
Search documents
Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
AI Engineer· 2026-08-22 17:00
Healthcare AI Deployment and Error Risks - Ambient scribes are deployed rapidly across approximately 33% (one-third) of United States medical practices and continue to grow[6] - Physician artificial intelligence adoption doubled over the preceding year, while adverse event reporting remains entirely untracked for most systems[6] - Approximately 5% (1 in 20) of real-world production notes contain errors serious enough to cause significant patient harm[5] - Nearly 20% (nearly 1 in 5) of all notes exhibit important omissions, and over 10% (more than 1 in 10) contain hallucinations[5] Evaluation System Limitations and Methodologies - Standard out-of-the-box frontier judges and pre-specified rubrics fail to catch critical contextual errors, passing roughly 20% (1 in 5) of notes that contain serious buried errors[24][44] - High-stakes evaluation requires capturing expert judgment and case-specific corrections rather than relying on static prompts or model weight retraining[33][34] - Effective evaluation operates through an evolving three-step loop of discovering real-world failure modes, capturing expert feedback, and calibrating outputs contextually[35][41]