← ObservatoryThe RecordFR-AI-0007
PROG-AI
FR-AI-0007

Autonomous AI Scientific Discovery — Novel, Correct, Independent

AI systems can autonomously conduct scientific research that produces novel, correct discoveries.

FragmentingVS-03·since 2026-09-05
Assessment trajectory
Fragmentingstate held · last assessed 2026-09-05
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
VS-03
Audit
Current from 2024-01-15 — present
VS-04
Replication
VS-05
Operation
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateFragmentingVS-03
Why this state?Governed assessment correction following LPR-001-D07. AS-001 and AS-002 remain visible append-only as historical judgements. IN-001 through IN-005 were corrected or bounded; no new evidence instance was admitted through LPR-001 and no verification-stage change was made.
Assessment summaryLPR-001-D07 materially narrows the historical basis of AS-001 and the GNoME/FunSearch shorthand inherited by AS-002 without rewriting those assessments. Corrected IN-002 separates GNoME's large-scale computational materials discovery from the distinct A-Lab autonomous-synthesis experiment; together they support bounded generation and experimental autonomy, but not a single end-to-end system that autonomously selected a problem and experimentally confirmed 736 discoveries. Corrected IN-004 still supplies strong evidence of novel, verifiable mathematical and algorithmic results from an AI-centred search process, but within a human-specified evaluator, program skeleton and problem. IN-001 is now correctly bounded to blind high-accuracy prediction, and IN-003 demonstrates an automated research loop evaluated by an automated reviewer rather than external conference acceptance. IN-005 is withdrawn from the current evidential basis because its institutional-restructuring provenance could not be reconstructed. The later evidence remains mixed: IN-007 and IN-008 expose provenance and scientific-judgment failure modes, while IN-009 provides supportive preprint evidence for autonomous research direction within a human-defined shared goal. FRAGMENTING / VS-03 is retained. The current record supports meaningful autonomous scientific work in bounded human-framed settings, but it does not yet converge on autonomous problem identification, robust scientific judgment, and independently verified novelty/correctness as a single general capability.
State entered2024-01-15
Last reaffirmed2026-09-05
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

BottleneckBN-001

"Autonomously conduct" lacks an agreed boundary. The claim requires autonomous research conduct, but the boundary between autonomous AI research and AI-assisted human research is contested. Current leading examples demonstrate substantial autonomous problem-solving, generation, experimentation, or research direction inside human-framed objectives. GNoME/A-Lab and FunSearch no longer support the stronger shorthand that the full autonomy component is satisfied; IN-009 extends autonomy toward research-direction choice while retaining a human-defined shared goal. Whether the claim requires only autonomous problem-solving within a supplied domain or also autonomous identification of the scientific problem remains a critical definitional gap. IN-008 further shows that scientific judgment after a problem is supplied is a separable autonomy constraint.

BottleneckBN-002

Novelty assessment is itself a research task. The claim requires that discoveries be novel, but establishing novelty requires surveying the accessible scientific literature — which is itself an incomplete and poorly indexed object. For fast-moving fields, a result that appears novel may have been anticipated in preprints, conference talks, or unpublished work. For large, old literatures, a result that appears novel may rediscover forgotten work. Novelty is not directly measurable from the discovery alone; it requires a comparison to the state of knowledge, which is itself uncertain. This is a measurement validity bottleneck of the same type as FR-BT-0002 BN-001: the measurement tool (literature survey) may not reliably track the thing it purports to measure (genuine novelty).

AttractorAT-001

Autonomous problem identification with verified novel correct results. The resolution path is a demonstration where an AI system identifies a previously unrecognised scientific problem, generates hypotheses about it, designs or conducts experiments, and produces results that are independently verified as correct and novel — without a human specifying the problem space. GNoME/A-Lab and FunSearch demonstrate important but bounded pieces of this path inside human-framed objectives; IN-009 adds evidence of research-direction choice within a human-defined shared goal. No current instance establishes the full attractor. The remaining gap includes both autonomous problem identification and sufficiently robust scientific judgment and validation once research is underway.

Assessment History
2024-01-15
Initial assessment — Fragmenting
The evidence is fragmenting across the three component claims. The correctness and novelty components are most strongly evidenced: GNoME (INST-002) and FunSearch (INST-004) both demonstrate AI systems producing results that are verified correct and independently novel in their domains. The autonomy component is more contested: in both cases, the research question was human-framed; the AI system discovered answers within a human-specified problem space rather than identifying the problem itself.
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-08-01
Reassessed, no change — Fragmenting
FRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.
AS-002 was issued following the operator-approved Post-Scout review of flag 2026-08-01-01. It is triggered by IN-008 (arXiv:2607.27191v1) and updates the current autonomy-boundary rationale without modifying AS-001, BN-001, OQ-001, AT-001, or any prior instance. The source is a primary preprint with two principal case studies and five total runs; its limits are preserved in IN-008.
2026-09-05
Reassessed, no change — Fragmenting
LPR-001-D07 materially narrows the historical basis of AS-001 and the GNoME/FunSearch shorthand inherited by AS-002 without rewriting those assessments. Corrected IN-002 separates GNoME's large-scale computational materials discovery from the distinct A-Lab autonomous-synthesis experiment; together they support bounded generation and experimental autonomy, but not a single end-to-end system that autonomously selected a problem and experimentally confirmed 736 discoveries. Corrected IN-004 still supplies strong evidence of novel, verifiable mathematical and algorithmic results from an AI-centred search process, but within a human-specified evaluator, program skeleton and problem. IN-001 is now correctly bounded to blind high-accuracy prediction, and IN-003 demonstrates an automated research loop evaluated by an automated reviewer rather than external conference acceptance. IN-005 is withdrawn from the current evidential basis because its institutional-restructuring provenance could not be reconstructed. The later evidence remains mixed: IN-007 and IN-008 expose provenance and scientific-judgment failure modes, while IN-009 provides supportive preprint evidence for autonomous research direction within a human-defined shared goal. FRAGMENTING / VS-03 is retained. The current record supports meaningful autonomous scientific work in bounded human-framed settings, but it does not yet converge on autonomous problem identification, robust scientific judgment, and independently verified novelty/correctness as a single general capability.
Governed assessment correction following LPR-001-D07. AS-001 and AS-002 remain visible append-only as historical judgements. IN-001 through IN-005 were corrected or bounded; no new evidence instance was admitted through LPR-001 and no verification-stage change was made.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

1955–90
Early AI discovery systems. DENDRAL (1965) and AM (1976) demonstrate early AI systems generating hypotheses in chemistry and mathematics. The claim's aspirational form is established; the capability is far from practical demonstration.
2020–22
AlphaFold2 establishes blind, near-experimental-accuracy protein-structure prediction as a major AI-for-science capability. The source supports autonomous execution of a human-defined prediction task, not autonomous selection of scientific questions or a full discovery loop.
2023
GNoME expands computational materials exploration while the distinct A-Lab platform demonstrates autonomous synthesis of human-selected targets; FunSearch produces novel cap-set constructions and improved bin-packing heuristics inside human-specified problem definitions. Novel generation and experimental or mathematical verification strengthen, but the evidence remains modular rather than one end-to-end autonomous scientist.
2024
The AI Scientist demonstrates an automated idea-to-paper research loop in machine learning, with evaluation by an automated reviewer rather than independent conference acceptance. The legacy institutional-restructuring bundle cannot be confidently sourced and is withdrawn from the current evidential basis. The claim remains FRAGMENTING as autonomy, scientific judgment, novelty and correctness diverge in evidential strength.
2026
New agent evidence deepens both sides of the record: data-poisoning and open-ended-research evaluations expose correctness and scientific-judgment failure modes, while the Station reports autonomous research-direction choice and novel mathematical results within a human-defined shared goal. The evidence moves beyond execution alone without yet demonstrating autonomous problem identification plus independently verified correctness and novelty as a general capability.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

Does "autonomously conduct scientific research" require autonomous problem identification, or is autonomous problem-solving within human-framed domains sufficient? BN-001 cannot close until this is resolved. The claim's satisfaction hangs on this distinction.

Raised 2024-01-15
OQ-002

The legacy IN-005 institutional-restructuring bundle could not be source-verified under LPR-001-D07 and is no longer part of the current evidential basis. The earlier question of whether institutional reorganisation constitutes a distinct anticipatory-act type therefore remains ungrounded for this record and should not be elevated from FR-AI-0007 unless a separately governed review establishes a source-faithful institutional evidence set.

Raised 2024-01-15
OQ-003

BN-002 (novelty assessment as a measurement validity bottleneck) is structurally similar to FR-BT-0002 BN-001 (biological age measurement validity). Both are cases where the measurement tool may not reliably track the thing it purports to measure. Two occurrences of this specific bottleneck structure across two programmes. Has measurement validity as a distinct resistance/bottleneck type now reached watchlist elevation?

Raised 2024-01-15
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0172026-09-05assessment_correctionAS-002AS-003
M-0162026-09-05provenance_correctionLPR-001-D07 discrepancies_foundLEGACY-INSTANCES-CORRECTED
M-0152026-09-05provenance_reviewLPR-001-D07
M-0142026-08-29instance_appendedIN-008IN-009
M-0132026-08-01assessment_issuedAS-001AS-002
M-0122026-08-01instance_appendedIN-007IN-008
M-0112026-07-17instance_appendedIN-007
M-0102026-07-14vector_correctedneutral--constrained-autonomy-boundary-untouchedNEUTRAL
M-0092026-07-14instance_appendedIN-006
M-0082026-07-09reference_correctedREFERENCE-CORRECTED
M-0072026-07-09description_restoredDESCRIPTION-RESTORED
M-0062026-07-09description_reorderedDESCRIPTION-REORDERED
M-0052024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
9 instances on recordShow sources ↓Hide ↑
IN-001AlphaFold2 — highly accurate blind protein-structure prediction1. Jumper, J. et al. (2021), Highly accurate protein structure prediction with AlphaFold, Nature 596, 583–589. · Abstract; CASP14 blind assessment; accuracy resultspartial
IN-002GNoME and A-Lab — computational materials discovery and autonomous synthesis are distinct results1. Merchant, A. et al. (2023), Scaling deep learning for materials discovery, Nature 624, 80–85. · Abstract; 2.2 million stable structures; 381,000 new convex-hull entries; 736 independently experimentally verified structures2. Szymanski, N. J. et al. (2023), An autonomous laboratory for the accelerated synthesis of inorganic materials, Nature 624, 86–91. · Abstract; 36 of 57 targets realised; autonomous synthesis workflow; target-selection limitationssupportive
IN-003The AI Scientist — automated research loop with automated evaluation1. Lu, C. et al. (2024), The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, arXiv:2408.06292. · Abstract; automated idea-to-paper workflow; simulated review; automated-reviewer acceptance thresholdpartial
IN-004FunSearch — novel cap-set constructions and improved bin-packing heuristics1. Romera-Paredes, B. et al. (2023), Mathematical discoveries from program search with large language models, Nature 625, 468–475. · Abstract; cap-set constructions; bin-packing heuristics; FunSearch specification and evaluatorsupportive
IN-005Legacy institutional-restructuring attribution — provenance unresolvedpartial
IN-006Structured Concept Evolution — LLM-driven discovery of qLDPC code families1. Liu, Z. & Marquardt, F. (2026), Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution, arXiv:2606.24808v1. · Abstract; structured concept evolution framework; code-capacity evaluationNEUTRAL
IN-007Distributed Denial of Science — indirect data poisoning of autonomous research agents1. Gyevnár, B., Kasirzadeh, A. & Shah, N. B. (2026), Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud, arXiv:2607.10712v1. · Abstract; five topics, three frontier systems, 450 runs; poisoning and mitigation resultsCONTESTING
IN-008Open-ended AI research case studies — engineering competence without successful scientific judgment1. Kirgis, P. et al. (2026), Can AI agents conduct open-ended AI research? Early evidence from two case studies, arXiv:2607.27191v1. · Abstract; shadow-evaluation design; two unpublished NeurIPS 2026 research questions; robustness checkCONTESTING
IN-009The Station — autonomous mathematical discovery in an open-world multi-agent environment1. Chung, S., Du, W. & Wesley, W. J. (2026), Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment, arXiv:2608.23691v1. · Abstract; open-world multi-agent environment; reported novel results; released dialogues, proofs and verification codeSUPPORTIVE