AI-assisted medical diagnosis achieves specialist-level accuracy on defined imaging tasks.
Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.
Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.
Research-dataset performance is an imperfect proxy for clinical deployment performance. The claim is defined around accuracy on imaging tasks, but the evidential meaning of an AUC, sensitivity, specificity or F1 result depends on the population, acquisition environment, reference standard and workflow in which it is measured. Cross-site studies show that performance can change under distribution shift, while prospective trials such as MASAI show that some AI-supported workflows can succeed clinically. The bottleneck is therefore not whether prospective clinical success is possible — MASAI now demonstrates that it is — but how reliably performance transfers across imaging domains, populations, acquisition environments and implementations.
Distribution shift and confounding can degrade transfer. Medical-imaging models may exploit correlations associated with acquisition site, equipment, population or workflow that do not remain stable elsewhere. Zech et al. directly demonstrate site-associated confounding in chest radiography, and cross-dataset work shows domain shift can reduce performance. This is a documented resistance mechanism, not a claim that every deep-learning imaging system necessarily fails for the same reason.
Replicated prospective multi-site clinical performance across distinct imaging domains, populations and implementations. MASAI now provides mature randomised evidence that an AI-supported mammography workflow can maintain or improve clinically meaningful accuracy in population screening. Resolution of the record's deployment-depth question requires comparable evidence beyond a single domain and programme: stable performance across diverse sites, patient populations, acquisition systems and clinical workflows, with patient-relevant outcomes and independent replication.
Historical narrative recorded for this claim. It does not override the current State Warrant.
Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.
BN-001 is the first measurement validity bottleneck in PROG-AI. RN-005 was developed from PROG-BT and PROG-AI (FR-AI-0007) evidence. Does its appearance here in an applied deployment claim — rather than a frontier capability claim — suggest that measurement validity is a broader AI phenomenon than previously established, or does it reflect a property specific to the benchmark-deployment gap in applied AI?
Raised 2024-01-15The deployment generalisation gap is a depth question structurally different from foundational uncertainty. Is this a new depth-layer category within PROG-AI, or is it the same surface/depth inversion described differently? If AI systems cannot reliably generalise their demonstrated capabilities to deployment environments, that is a depth failure even for surface claims.
Raised 2024-01-15MASAI now demonstrates mature prospective success in one population-screening workflow. What breadth of replication is required to close BN-001: independent multi-site replication within mammography, successful transfer across populations and acquisition systems, or comparable prospective evidence across several imaging domains?
Raised 2024-01-15| Mutation | Date | Field | Prior value | Current value |
|---|---|---|---|---|
| M-011 | 2026-09-06 | assessment_issued | AS-002 | AS-003 |
| M-010 | 2026-09-06 | instance_added | IN-005 | IN-006 |
| M-009 | 2026-09-06 | assessment_and_dependencies_corrected | AS-001 / legacy BN-RM-AT-lineage-OQ wording | AS-002 / corrected dependencies |
| M-008 | 2026-09-06 | provenance_correction | LPR-001-D08 discrepancies_found | LPR-001-D08 discrepancies_corrected |
| M-007 | 2026-09-06 | provenance_review | — | LPR-001-D08 |
| M-006 | 2026-07-09 | description_reordered | — | DESCRIPTION-REORDERED |
| M-005 | 2024-01-15 | diagnosis_held | — | DIAGNOSIS-HELD |
| M-004 | 2024-01-15 | mechanisms_recorded | — | MECHANISMS-RECORDED |
| M-003 | 2024-01-15 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-002 | 2024-01-15 | instances_logged | — | INSTANCES-LOGGED |
| M-001 | 2024-01-15 | record_created | — | RECORD-CREATED |