← ObservatoryThe RecordFR-AI-0008
PROG-AI
FR-AI-0008

AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined Tasks

AI-assisted medical diagnosis achieves specialist-level accuracy on defined imaging tasks.

FragmentingVS-03·since 2026-09-06
Assessment trajectory
Fragmentingstate held · last assessed 2026-09-06
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
VS-03
Audit
Current from 2026-09-06 — present
VS-04
Replication
VS-05
Operation
First recorded 2024-01-15
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateFragmentingVS-03
Why this state?Bounded Record Review of Gommers et al. (2026), The Lancet, DOI 10.1016/S0140-6736(25)02464-X. IN-006 is admitted as new evidence with primary provenance. Historical assessments remain append-only; no legacy instance is rewritten through this Record Review.
Assessment summaryNormal Record Review admits IN-006, the completed MASAI trial's primary interval-cancer analysis. The result materially strengthens the prospective-deployment side of the record: AI-supported mammography was non-inferior for interval-cancer rate, had significantly higher sensitivity, the same specificity, and fewer interval cancers with several unfavourable characteristics, while earlier MASAI analyses showed increased cancer detection and markedly reduced reading workload. This means the benchmark-to-deployment gap is no longer represented only by heterogeneous or preliminary prospective evidence; one large randomised population-screening programme now provides mature clinical evidence of maintained or improved diagnostic performance. The evidence still does not converge across medical imaging as a whole. IN-003 documents genuine cross-site and domain-shift failures, IN-005 provides no verified foundation-model resolution, and MASAI remains a domain-specific workflow in a Swedish screening context. FRAGMENTING / VS-03 is therefore retained, but the positive prospective pole is materially stronger and BN-001 is narrowed from a general benchmark-versus-deployment uncertainty to a transferability question across domains, populations and implementations.
State entered2024-01-15
Last reaffirmed2026-09-06
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

BOTTLENECK — MEASUREMENT VALIDITY (RN-005)BN-001

Research-dataset performance is an imperfect proxy for clinical deployment performance. The claim is defined around accuracy on imaging tasks, but the evidential meaning of an AUC, sensitivity, specificity or F1 result depends on the population, acquisition environment, reference standard and workflow in which it is measured. Cross-site studies show that performance can change under distribution shift, while prospective trials such as MASAI show that some AI-supported workflows can succeed clinically. The bottleneck is therefore not whether prospective clinical success is possible — MASAI now demonstrates that it is — but how reliably performance transfers across imaging domains, populations, acquisition environments and implementations.

Resistance MechanismRM-001

Distribution shift and confounding can degrade transfer. Medical-imaging models may exploit correlations associated with acquisition site, equipment, population or workflow that do not remain stable elsewhere. Zech et al. directly demonstrate site-associated confounding in chest radiography, and cross-dataset work shows domain shift can reduce performance. This is a documented resistance mechanism, not a claim that every deep-learning imaging system necessarily fails for the same reason.

AttractorAT-001

Replicated prospective multi-site clinical performance across distinct imaging domains, populations and implementations. MASAI now provides mature randomised evidence that an AI-supported mammography workflow can maintain or improve clinically meaningful accuracy in population screening. Resolution of the record's deployment-depth question requires comparable evidence beyond a single domain and programme: stable performance across diverse sites, patient populations, acquisition systems and clinical workflows, with patient-relevant outcomes and independent replication.

Assessment History
2024-01-15
Initial assessment — Fragmenting
The claim's surface assertion — specialist-level accuracy on defined imaging tasks — is confirmed on curated research datasets across multiple imaging domains and by regulatory validation in prospective settings for specific cleared devices. The surface layer is advancing: AI medical imaging achieves specialist-level performance on well-defined tasks under controlled conditions. The surface claim is in ESCALATING territory. The claim fragments at the depth layer — specifically, at the boundary between research-dataset accuracy and real-world clinical deployment. Systematic deployment-gap studies (INST-003) document that accuracy measured on curated, single-site datasets does not reliably generalise across scanners, acquisition protocols, or patient demographics, and prospective trials (INST-004) show a heterogeneous picture — some deployed systems retain specialist-level accuracy, others do not. The pressure state is FRAGMENTING: the surface claim is confirmed and advancing, but the depth question — whether research-dataset accuracy is a valid proxy for clinical deployment accuracy — remains open (BN-001), pending further validation of the foundation-model generalisation trend (INST-005).
Verification Stage: VS-05 after ratified review (stored code VS-03 preserved).
2026-09-06
Reassessed, no change — Fragmenting
LPR-001-D08 materially narrows the evidential basis without reversing the record. IN-001 still supports specialist-level performance on bounded research tasks across dermatology, diabetic-retinopathy screening and chest-radiograph pneumonia detection. IN-002 establishes device-specific regulatory authorisation, but the legacy inference that FDA clearance itself demonstrates prospective specialist-level clinical performance is withdrawn. IN-003 supports a real external-validation and evidence-quality problem, but not a universal causal claim that deep learning necessarily learns only dataset-specific features. IN-004 now supplies the strongest prospective clinical evidence in the record: the randomised MASAI mammography trial shows improved cancer detection with substantially reduced reading workload and no significant increase in false positives. IN-005 no longer supports the proposition that foundation models have already reduced medical-imaging distribution shift. The evidence therefore remains fragmented between strong bounded-task performance and heterogeneous evidence about transfer into clinical environments. FRAGMENTING / VS-03 is retained.
Append-only correction following LPR-001-D08. AS-001 is preserved as historical assessment. This assessment removes reliance on the unsupported prospective-clearance and foundation-model-generalisation premises while retaining the benchmark-to-deployment boundary on narrower evidence. The 2026 final MASAI interval-cancer analysis is not admitted here and remains a separate Record Review candidate.
2026-09-06
Reassessed, no change — Fragmenting
Normal Record Review admits IN-006, the completed MASAI trial's primary interval-cancer analysis. The result materially strengthens the prospective-deployment side of the record: AI-supported mammography was non-inferior for interval-cancer rate, had significantly higher sensitivity, the same specificity, and fewer interval cancers with several unfavourable characteristics, while earlier MASAI analyses showed increased cancer detection and markedly reduced reading workload. This means the benchmark-to-deployment gap is no longer represented only by heterogeneous or preliminary prospective evidence; one large randomised population-screening programme now provides mature clinical evidence of maintained or improved diagnostic performance. The evidence still does not converge across medical imaging as a whole. IN-003 documents genuine cross-site and domain-shift failures, IN-005 provides no verified foundation-model resolution, and MASAI remains a domain-specific workflow in a Swedish screening context. FRAGMENTING / VS-03 is therefore retained, but the positive prospective pole is materially stronger and BN-001 is narrowed from a general benchmark-versus-deployment uncertainty to a transferability question across domains, populations and implementations.
Bounded Record Review of Gommers et al. (2026), The Lancet, DOI 10.1016/S0140-6736(25)02464-X. IN-006 is admitted as new evidence with primary provenance. Historical assessments remain append-only; no legacy instance is rewritten through this Record Review.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

2016–17
Landmark bounded-task studies. Dermatology, diabetic-retinopathy and chest-radiograph studies establish specialist-level or specialist-comparable performance under defined research evaluations. They do not by themselves establish deployment generalisation.
2018–20
Regulatory adoption and generalisation scrutiny develop in parallel. IDx-DR and Viz.ai demonstrate device-specific FDA authorisation, while Zech, Pooch and the Nagendran review show why internal performance and regulatory status should not be conflated with uniform real-world generalisation.
2021–25
Prospective clinical evidence strengthens unevenly. MASAI provides randomised evidence that an AI-supported mammography workflow can increase cancer detection while reducing reading workload without a significant false-positive increase. The broader cross-domain deployment question remains open.
2023–24
Foundation-model generalisation remains unestablished in this legacy evidence set. The former IN-005 bundle is withdrawn from the current evidential basis because its cited sources do not establish improved cross-site medical-imaging generalisation.
2026
The completed MASAI primary analysis strengthens the prospective clinical pole. AI-supported screening is non-inferior for interval-cancer rate, significantly more sensitive, and equally specific to standard double reading in the trial population. The record's unresolved question shifts toward whether similarly mature performance transfers across other imaging domains and deployment environments.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

BN-001 is the first measurement validity bottleneck in PROG-AI. RN-005 was developed from PROG-BT and PROG-AI (FR-AI-0007) evidence. Does its appearance here in an applied deployment claim — rather than a frontier capability claim — suggest that measurement validity is a broader AI phenomenon than previously established, or does it reflect a property specific to the benchmark-deployment gap in applied AI?

Raised 2024-01-15
OQ-002

The deployment generalisation gap is a depth question structurally different from foundational uncertainty. Is this a new depth-layer category within PROG-AI, or is it the same surface/depth inversion described differently? If AI systems cannot reliably generalise their demonstrated capabilities to deployment environments, that is a depth failure even for surface claims.

Raised 2024-01-15
OQ-003

MASAI now demonstrates mature prospective success in one population-screening workflow. What breadth of replication is required to close BN-001: independent multi-site replication within mammography, successful transfer across populations and acquisition systems, or comparable prospective evidence across several imaging domains?

Raised 2024-01-15
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0112026-09-06assessment_issuedAS-002AS-003
M-0102026-09-06instance_addedIN-005IN-006
M-0092026-09-06assessment_and_dependencies_correctedAS-001 / legacy BN-RM-AT-lineage-OQ wordingAS-002 / corrected dependencies
M-0082026-09-06provenance_correctionLPR-001-D08 discrepancies_foundLPR-001-D08 discrepancies_corrected
M-0072026-09-06provenance_reviewLPR-001-D08
M-0062026-07-09description_reorderedDESCRIPTION-REORDERED
M-0052024-01-15diagnosis_heldDIAGNOSIS-HELD
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
6 instances on recordShow sources ↓Hide ↑
IN-001Landmark studies — specialist-level performance on bounded imaging benchmarks1. Esteva, A. et al. (2017), Dermatologist-level classification of skin cancer with deep neural networks, Nature 542, 115–118. · Abstract; training set; dermatologist comparison2. Gulshan, V. et al. (2016), Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs, JAMA 316(22), 2402–2410. · Results; EyePACS-1 and Messidor-2 validation; sensitivity and specificity3. Rajpurkar, P. et al. (2017), CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning, arXiv:1711.05225. · Abstract; ChestX-ray14; four-radiologist comparison; F1 metricsupportive
IN-002FDA authorisations — clinical-use evidence with device-specific validation1. U.S. FDA (2018), De Novo classification record DEN180001 — IDx-DR. · Decision date 11 April 2018; diabetic retinopathy detection device2. U.S. FDA (2018), FDA permits marketing of clinical decision support software for alerting providers of a potential stroke in patients. · Intended use; retrospective 300-CT study; real-world notification evidence; diagnostic limitationsupportive
IN-003External validation and evidence-quality limits — deployment generalisation remains conditional1. Zech, J. R. et al. (2018), Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study, PLOS Medicine 15(11):e1002683. · Abstract; three-hospital cross-site evaluation; confounding/site detection2. Pooch, E. H. P., Ballester, P. L. & Barros, R. C. (2019), Can we trust deep learning models diagnosis? The impact of domain shift in chest radiograph classification, arXiv:1909.01940. · Abstract; cross-dataset domain-shift evaluation3. Nagendran, M. et al. (2020), Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies, BMJ 368:m689. · Methods; principal findings; prospective and randomised-study limitationscontesting
IN-004MASAI — prospective randomised AI-supported mammography screening1. Dembrower, K. et al. (2025), Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI), The Lancet Digital Health. · Randomised screening design; cancer detection; false positives; 44.2% workload reductionsupportive
IN-005Legacy foundation-model generalisation attribution — unsupported as writtenpartial
IN-006MASAI final interval-cancer analysis — favourable prospective clinical performance1. Gommers, J. et al. (2026), Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study, The Lancet 407(10527), 505–514. · Primary interval-cancer outcome; sensitivity; specificity; interval-cancer characteristicssupportive