Faultline Observatory

Evidence Trajectories

Follow how the Observatory's judgement of technology claims has changed as evidence accumulated.

Each line represents one Frontier Record. Select a reading to bring related trajectories forward; the full archive remains present.

How to read this page
  1. Read across - time moves from left to right.
  2. Read vertically - position indicates verification depth at that moment.
  3. Single-assessment records - one documented assessment is shown as one historical point, not as a synthetic line to Today.
  4. Read provenance - dashed rings mark single-assessment stages touched by the legacy review; finer dotted rings mark historically unverified assignments.
  5. Read the register - the chart on the left shows documentary history; the separate Current Register on the right groups records by their ratified Verification Stage today.
  6. Open the record - every trajectory can be traced back to its documentary history.

Verification-stage provenance notice: the legacy review is complete. Historical codes remain preserved. For single-assessment records, the historical chart now shows the recorded stage while the Current Register shows the ratified present interpretation. Assignments that could not be reliably reconstructed remain visibly marked historically unverified.

Evidence TrajectoriesEvidence Trajectories arranges documented assessments across four evidence-history phases. Horizontal spacing represents documentary sequence rather than equal calendar time. Single-assessment records are shown as one historical point rather than a synthetic full-width line. The chart on the left shows documentary history; the separate Current Register on the right groups records by their ratified verification stage today.FR-QE-0001, 2026-06-11, recorded stage VS-03, Fragmenting.FR-QE-0001, 2026-07-22, recorded stage VS-04, Stabilising.FR-QE-0002, 2024-01-15, recorded stage VS-03, Fragmenting.FR-QE-0002, 2026-07-26, recorded stage VS-03, Fragmenting.FR-QE-0003, 2024-01-15, recorded stage VS-02, Escalating.FR-QE-0003, 2026-06-28, recorded stage VS-03, Escalating.FR-QE-0004, 2024-01-15, recorded stage VS-04, Resolving.FR-QE-0004, 2026-08-17, recorded stage VS-04, Resolving.FR-QE-0005, 2024-01-15, recorded stage VS-02, Escalating.FR-QE-0005, 2026-06-29, recorded stage VS-03, Escalating.FR-QE-0006, 2024-01-15, recorded stage VS-02, Escalating.FR-QE-0006, 2026-08-17, recorded stage VS-02, Escalating.FR-QE-0007, 2024-01-15, recorded stage VS-03, Fragmenting.FR-QE-0007, 2026-08-28, recorded stage VS-03, Fragmenting.FR-QE-0008, 2024-01-15, recorded stage VS-04, Resolving. Legacy review marked this stage historically unverified with low confidence on 2026-07-23.FR-AI-0001, 2024-01-15, recorded stage VS-02, Escalating.FR-AI-0001, 2026-06-27, recorded stage VS-03, Escalating.FR-AI-0002, 2024-01-15, recorded stage VS-02, Escalating.FR-AI-0002, 2026-08-17, recorded stage VS-02, Escalating.FR-AI-0003, 2024-01-15, recorded stage VS-03, Fragmenting.FR-AI-0003, 2026-07-14, recorded stage VS-03, Fragmenting.FR-AI-0003, 2026-09-01, recorded stage VS-03, Fragmenting.FR-AI-0004, 2024-01-15, recorded stage VS-03, Fragmenting.FR-AI-0004, 2026-06-29, recorded stage VS-03, Fragmenting.FR-AI-0004, 2026-09-02, recorded stage VS-03, Fragmenting.FR-AI-0005, 2024-01-15, recorded stage VS-03, Fragmenting.FR-AI-0005, 2026-06-29, recorded stage VS-03, Fragmenting.FR-AI-0005, 2026-09-03, recorded stage VS-03, Fragmenting.FR-AI-0006, 2024-01-15, recorded stage VS-03, Fragmenting.FR-AI-0006, 2026-08-29, recorded stage VS-03, Fragmenting.FR-AI-0006, 2026-09-04, recorded stage VS-03, Fragmenting.FR-AI-0006, 2026-09-04, recorded stage VS-03, Fragmenting.FR-AI-0007, 2024-01-15, recorded stage VS-03, Fragmenting.FR-AI-0007, 2026-08-01, recorded stage VS-03, Fragmenting.FR-AI-0007, 2026-09-05, recorded stage VS-03, Fragmenting.FR-AI-0008, 2024-01-15, recorded stage VS-05, Fragmenting.FR-AI-0008, 2026-09-06, recorded stage VS-03, Fragmenting.FR-AI-0008, 2026-09-06, recorded stage VS-03, Fragmenting.FR-AI-0009, 2026-08-19, recorded stage VS-02, Escalating.FR-AM-0001, Mar 1989, recorded stage VS-01, Emerging.FR-AM-0001, Apr–Nov 1989, recorded stage VS-04, Fragmenting.FR-AM-0001, 2004 or earlier, recorded stage VS-04, Collapsed.FR-AM-0001, 2026-09-08, recorded stage VS-05, Collapsed.FR-AM-0001, 2026-09-08, recorded stage VS-05, Collapsed.FR-AM-0002, 2024-01-15, recorded stage VS-04, Fragmenting.FR-AM-0002, 2026-09-09, recorded stage VS-03, Fragmenting.FR-AM-0003, 2024-01-15, recorded stage VS-03, Fragmenting. Legacy review re-affirmed this stage with high confidence on 2026-07-23.FR-AM-0004, 2024-01-15, recorded stage VS-03, Escalating.FR-AM-0004, 2026-06-29, recorded stage VS-03, Escalating.FR-AM-0005, 2024-01-15, recorded stage VS-04, Collapsed.FR-AM-0005, 2026-06-29, recorded stage VS-04, Collapsed.FR-AM-0006, 2024-01-15, recorded stage VS-03, Escalating.FR-AM-0006, 2026-06-27, recorded stage VS-03, Escalating.FR-AM-0006, 2026-08-29, recorded stage VS-03, Escalating.FR-AM-0006, 2026-08-29, recorded stage VS-02, Escalating.FR-AM-0007, 2026-08-25, recorded stage VS-03, Escalating.FR-BT-0001, 2024-01-15, recorded stage VS-02, Escalating. Legacy review reconstructed the current stage from VS-02 to VS-03 with medium-high confidence on 2026-07-23.FR-BT-0002, 2024-01-15, recorded stage VS-03, Escalating.FR-BT-0002, 2026-06-29, recorded stage VS-03, Escalating.FR-BT-0002, 2026-08-29, recorded stage VS-02, Escalating.FR-BT-0003, 2024-01-15, recorded stage VS-03, Fragmenting. Legacy review reconstructed the current stage from VS-03 to VS-04 with medium-high confidence on 2026-07-23.FR-BT-0004, 2024-01-15, recorded stage VS-03, Fragmenting.FR-BT-0004, 2026-06-27, recorded stage VS-04, Fragmenting.FR-BT-0005, 2026-08-21, recorded stage VS-03, Escalating.OperationVS-05Demonstrated operationFR-AM-0001: Cold Fusion — Room-Temperature Nuclear Fusion in Electrochemical Cells. Current verification stage VS-05 Operation.FR-AM-0001Cold FusionReplicationVS-04FR-AM-0005: Room-Temperature Superconductivity — Reproducibility Under Laboratory Conditions. Current verification stage VS-04 Replication.FR-AM-0005Room-Temperature SuperconductivityFR-BT-0003: Biological Age Biomarker Panels — Predictive Validity for Age-Related Decline. Current verification stage VS-04 Replication. Legacy review reconstructed the current stage from VS-03 to VS-04 with medium-high confidence on 2026-07-23.FR-BT-0003Biological Age Biomarker PanelsFR-BT-0004: Liquid Biopsy — Early Cancer Detection Before Conventional Diagnosis. Current verification stage VS-04 Replication.FR-BT-0004Liquid BiopsyFR-QE-0001: Google Quantum Advantage (Sycamore) — Computational Supremacy via Random Circuit Sampling. Current verification stage VS-04 Replication.FR-QE-0001Google Quantum Advantage (Sycamore)FR-QE-0004: Below-Threshold Quantum Error Correction — Scalable Architecture. Current verification stage VS-04 Replication.FR-QE-0004Below-Threshold Quantum Error CorrectionFR-QE-0008: Quantum Error Correction Scaling — Logical Rate Suppression Under Physical Overhead. Current verification stage VS-04 Replication. Legacy review marked this stage historically unverified with low confidence on 2026-07-23.FR-QE-0008Quantum Error Correction ScalingAuditVS-03FR-AI-0001: LLM Multi-Step Reasoning — Generalisation Beyond Training. Current verification stage VS-03 Audit.FR-AI-0001LLM Multi-Step ReasoningFR-AI-0003: RLHF Preference Generalisation — Behaviour Beyond Training Distribution. Current verification stage VS-03 Audit.FR-AI-0003RLHF Preference GeneralisationFR-AI-0004: Scaling Laws — Emergent Performance on Unseen Tasks. Current verification stage VS-03 Audit.FR-AI-0004Scaling LawsFR-AI-0005: AGI Through Scaling — LLM Architecture as the Path to General Intelligence. Current verification stage VS-03 Audit.FR-AI-0005AGI Through ScalingFR-AI-0006: Scaling Mechanism Coherence — Continuity Across Model Sizes. Current verification stage VS-03 Audit.FR-AI-0006Scaling Mechanism CoherenceFR-AI-0007: Autonomous AI Scientific Discovery — Novel, Correct, Independent. Current verification stage VS-03 Audit.FR-AI-0007Autonomous AI Scientific DiscoveryFR-AI-0008: AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined Tasks. Current verification stage VS-03 Audit.FR-AI-0008AI Medical Imaging DiagnosisFR-AM-0002: Anomalous Excess Heat — Electrochemical Cells Beyond Conventional Chemistry. Current verification stage VS-03 Audit.FR-AM-0002Anomalous Excess HeatFR-AM-0003: Cuprate Superconductivity — Mechanism Identification. Current verification stage VS-03 Audit. Legacy review re-affirmed this stage with high confidence on 2026-07-23.FR-AM-0003Cuprate SuperconductivityFR-AM-0004: Commercial Fusion Power — Net Electricity at Grid Scale. Current verification stage VS-03 Audit.FR-AM-0004Commercial Fusion PowerFR-AM-0007: Pressure-Quenched Superconductivity — Retention of High-Pressure States at Ambient Pressure. Current verification stage VS-03 Audit.FR-AM-0007Pressure-Quenched SuperconductivityFR-BT-0001: Senolytic Therapies — Meaningful Human Healthspan Extension. Current verification stage VS-03 Audit. Legacy review reconstructed the current stage from VS-02 to VS-03 with medium-high confidence on 2026-07-23.FR-BT-0001Senolytic TherapiesFR-BT-0005: Gene-Edited Porcine Kidneys — Durable Human Renal Replacement. Current verification stage VS-03 Audit.FR-BT-0005Gene-Edited Porcine KidneysFR-QE-0002: D-Wave Quantum Annealing — Practical Computational Advantage. Current verification stage VS-03 Audit.FR-QE-0002D-Wave Quantum AnnealingFR-QE-0003: Fault-Tolerant Logical Qubits — Error Rate Scaling with Code Distance. Current verification stage VS-03 Audit.FR-QE-0003Fault-Tolerant Logical QubitsFR-QE-0005: Cryptographically Relevant Quantum Computing — RSA Factorisation. Current verification stage VS-03 Audit.FR-QE-0005Cryptographically Relevant Quantum ComputingFR-QE-0007: Practical Quantum Advantage — Performance Beyond Classical Computation on Relevant Problems. Current verification stage VS-03 Audit.FR-QE-0007Practical Quantum AdvantagePublishedVS-02FR-AI-0002: LLM Knowledge-Work Utility — Economically Valuable Task Performance. Current verification stage VS-02 Published.FR-AI-0002LLM Knowledge-Work UtilityFR-AI-0009: World Models — Physical Prediction and Transfer. Current verification stage VS-02 Published.FR-AI-0009World ModelsFR-AM-0006: Solid-State Batteries — Commercial Viability for Electric Vehicles. Current verification stage VS-02 Published.FR-AM-0006Solid-State BatteriesFR-BT-0002: Epigenetic Reprogramming — Biological Age Reversal Without Identity Loss. Current verification stage VS-02 Published.FR-BT-0002Epigenetic ReprogrammingFR-QE-0006: Fault-Tolerant Quantum Utility — Practically Useful Algorithms Beyond Classical Simulation. Current verification stage VS-02 Published.FR-QE-0006Fault-Tolerant Quantum UtilityAssertionVS-01No current records

The trajectories above are derived from the assessment histories below.

Trajectory records

SHOWING THE FULL ARCHIVE · NO READING IS PRESELECTED.

FR-QE-0001Google Quantum Advantage (Sycamore) — Computational Supremacy via Random Circuit SamplingA programmable quantum processor has demonstrated computational supremacy — performing a well-defined sampling task beyond the practical reach of any classical computer.QE · 2 assessments · 1 documented state change
Stabilisingsince 2026-07-22
Assessment history
  1. 2026-06-11Fragmenting

    The original 2019 quantum supremacy claim has undergone a complete evidence cycle. At announcement, the claim was precise and measurable: 200 seconds versus an estimated 10,000 classical years on a specific Random Circuit Sampling task. Within five years, classical simulation methods improved by orders of magnitude, culminating in the Zhao et al. (2024) result demonstrating classical performance exceeding the original quantum benchmark in both speed and fidelity. The original claim, as stated in 2019, has been effectively superseded. However, the research programme that produced the claim has not collapsed. Google's Willow processor (December 2024) reasserts the supremacy framing at a vastly larger scale (10²⁵ classical years) while simultaneously demonstrating below-threshold quantum error correction — a qualitatively different and more durable achievement. The evidence trajectory has therefore split: the narrow 2019 benchmark claim is contested to the point of supersession, while the broader programme claim (that quantum processors are advancing toward practical computational advantage) has arguably strengthened. This record exhibits the canonical pattern the FCIF subsequently formalised as Claim Migration: the original claim does not resolve cleanly (neither fully vindicated nor retracted) but instead evolves as the claimant shifts the evidential basis to a new formulation that inherits the original's ambition but rests on different technical foundations. The 2019 claim migrated from "supremacy via RCS on Sycamore" to "scalable error correction via surface codes on Willow." The Observatory records this as a Fragmenting state: the evidence does not converge on a single verdict because the claim itself has moved.

  2. 2026-07-22Stabilising

    Institutional verdict: UNRESOLVED AT THE ORIGINAL COMPARISON POINT; LATER SUPERSEDED. The admitted claim is a time-indexed comparative-performance claim concerning Sycamore's 2019 Random Circuit Sampling demonstration. The experiment and task performance are supported, but the constitutive claim that the task was beyond practical classical reach was not established at the original comparison point because IBM's contemporaneous days-scale analysis left that threshold unresolved. Later classical work, culminating in Zhao et al. (2024), reproduced and surpassed the benchmark comparator. That later result supersedes the demonstrated advantage without automatically proving that the time-indexed 2019 claim was false when made. Pressure State is STABILISING. The uncertainty is now bounded and durable rather than fragmenting: Willow and other successor claims are outside this claim's material commitments, and the original comparison can reopen only through evidence showing that the 2019 comparator was already unsound. Verification Stage is VS-04 — Replication because independent classical work progressed beyond audit to direct reproduction and eventual counter-performance of the claim's constitutive comparator. This is claim-level, adversarial replication; it does not assert independent reproduction of Sycamore hardware or favourable confirmation of the original advantage.

FR-QE-0002D-Wave Quantum Annealing — Practical Computational AdvantageQuantum annealing systems have demonstrated practical computational advantage over classical methods on commercially or scientifically relevant optimisation tasks.QE · 2 assessments · 0 documented state changes
Fragmentingsince 2026-07-26
Assessment history
  1. 2024-01-15Fragmenting

    The evidence trail for this claim is fragmented across distinct problem domains and claim interpretations. On commercially motivated optimisation tasks (scheduling, routing, combinatorial problems of practical scale), no published evidence has established durable advantage over state-of-the-art classical methods. The contested Denchev et al. (2016) result represents the strongest performance claim in this domain; it was substantially undermined by subsequent classical algorithm improvements and the benchmark's structural dependence on hardware-favourable problem instances (RM-002). In the separate domain of scientific simulation, the evidence is stronger: King et al. (2022, 2023) report computational advantage in simulating quantum magnetism, though critics dispute the comparison class used. The claim spans two domains accruing evidence asymmetrically and has not been decomposed into separate records (BN-001), and no agreed classical comparison class exists (BN-002). The pressure state is FRAGMENTING: the claim is not converging toward a single assessment but splitting along domain lines that may require separate evaluation.

  2. 2026-07-26Fragmenting

    The Pressure State is unchanged. Its governing rationale is not. The prior domain-split explanation (AS-001) is retired following the ratified Identity and Continuity Review, which found IDENTITY PRESERVED AS A COMPOUND CLAIM: a single recoverable kernel — quantum annealing, practical computational advantage, a classical comparator, optimisation-task class, commercial-or-scientific relevance — is engaged by evidence from both relevance routes. IN-004 is adjacent simulation evidence and does not bear on this claim. Among the remaining instances, IN-001 through IN-003 read negative-to-contested on the commercial route across eight years, and IN-005 provides a single, contemporaneously-grounded positive instance whose own comparator (quantum Monte Carlo) is disputed (BN-002). FRAGMENTING is warranted not because the claim splits along commercial or scientific lines, but because this same kernel-corrected evidence supports incompatible trajectory interpretations under unresolved competing meanings of 'practical advantage' — whether that standard requires real-world deployability or a rigorous demonstration of speedup on a well-posed instance (OQ-6, proposed). This is interpretive, not referential, non-convergence: no identity fracture, no decomposition, no admission-scope defect.

FR-QE-0003Fault-Tolerant Logical Qubits — Error Rate Scaling with Code DistanceFault-tolerant logical qubits can be demonstrated with logical error rates that improve as error-correcting code distance increases.QE · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-28
Assessment history
  1. 2024-01-15Escalating

    The claim describes a specific empirical signature: logical error rates improving as code distance increases. This signature has now been demonstrated. INST-003 (Google, 2023) was the first result to show simultaneous X and Z error suppression with increasing code distance, directly satisfying the claim's measurement criterion. INST-005 (Google Willow, 2024) extends this to below-threshold operation, showing that the improvement rate exceeds the overhead rate — the condition required for the result to be considered scalable rather than merely demonstrated at fixed size. INST-001 (2021) and INST-002 (2022) provided earlier partial evidence — single-error-type suppression and below-physical-error-rate operation, respectively — establishing the trajectory that INST-003 and INST-005 confirm more directly. The pressure state is ESCALATING: the claim's core empirical signature is demonstrated and strengthening, but the demonstrated code distances (up to 7) remain well below the distances required to confirm the behaviour holds at practically relevant scale (OQ-001).

  2. 2026-06-28Escalating

    The evidence gap is closed by IN-006. The Willow result is now treated as a verified, peer-reviewed below-threshold surface-code memory result rather than a general quantum-computing announcement. It materially strengthens the claim because logical error suppression improves with code distance and the larger memory exceeds break-even. The pressure state remains ESCALATING rather than RESOLVING because the record's own next decisive question — whether below-threshold scaling holds at d=11 and above — remains unanswered, and the demonstrated result is still a memory result rather than a full fault-tolerant computation pathway.

FR-QE-0004Below-Threshold Quantum Error Correction — Scalable ArchitectureQuantum error correction can reduce logical error rates below physical error rates in a scalable architecture.QE · 2 assessments · 0 documented state changes
Resolvingsince 2026-08-17
Assessment history
  1. 2024-01-15Resolving

    The claim has two components: below-physical-rate operation, and scalability of that operation. Both have been demonstrated. INST-002 established that below-physical-rate logical qubits are achievable in principle. INST-003 established that practically useful error rates are achievable on current hardware. INST-004 established that performance improves as the architecture scales — the defining signature of scalable below-threshold operation. No contesting evidence has been published against either component of the claim. The pressure state is RESOLVING: both elements the claim requires have been independently demonstrated and corroborate each other, though confirmation at the code distances required for practical fault-tolerant computation (d=11 and above) remains outstanding (OQ-001), and the claim's architecture-agnosticism has not yet been confirmed across a third hardware platform (OQ-002).

  2. 2026-08-17Resolving

    The record remains RESOLVING. The post-AS-001 evidence broadens the engineering case without crossing the remaining resolution boundary. IN-007 extends below-physical or breakeven behaviour to a qLDPC code on trapped-ion hardware; IN-008 demonstrates real-time recalibration that supports sustained error-corrected operation; and IN-009 shows that an essential entangling gate can preserve the favourable erasure-biased error hierarchy of superconducting dual-rail qubits. These are meaningful supportive developments across code family, sustained operation, and architecture. They do not, individually or together, demonstrate logical error suppression continuing at code distances d=11 and above, which remains OQ-001 and the record's decisive attractor. The contested Microsoft topological result remains corroborating rather than foundational. Pressure State therefore remains RESOLVING and Verification Stage remains VS-04; all three open questions remain live.

FR-QE-0005Cryptographically Relevant Quantum Computing — RSA FactorisationA quantum computer can factor commercially relevant RSA cryptographic keys faster than any classical computer.QE · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-29
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No quantum computer has factored a commercially relevant RSA key. The most credible direct attempt (INST-005) failed. The engineering gap between current capability and the Gidney-Ekerå resource estimate remains approximately three to four orders of magnitude in physical qubit count, with additional requirements for error rates, connectivity, and operational duration not yet demonstrated at any scale approaching relevance. The pressure state is ESCALATING rather than EMERGING because the substrate advances documented in FR-QE-0003 and FR-QE-0004 (INST-003) show the underlying error-correction engineering progressing on a credible trajectory, even though the gap to the resource requirement remains enormous. Institutional behaviour — NIST's finalisation of post-quantum cryptography standards (INST-004) — reflects institutional acceptance that the risk is credible enough to justify migration, adding pressure to the claim's trajectory independent of any direct technical progress toward satisfaction.

  2. 2026-06-29Escalating

    No threshold has been crossed since AS-001 — no factorisation of a commercially relevant key has occurred, and none is closer to occurring in any demonstrated sense. What has moved is the resource-estimate trajectory underlying OQ-001. Gidney (Google, May 2025) reduced the estimated physical-qubit requirement for RSA-2048 factorisation from the Gidney-Ekerå (2021) figure of ~20 million to under 1 million, under comparable fault-tolerance assumptions — roughly a 20-fold reduction achieved through improved algorithmic and error-correction engineering rather than any experimental demonstration. A 2026 proposal using QLDPC codes (an architecture distinct from the surface codes assumed in both prior estimates) suggests a further reduction toward ~100,000 physical qubits, though this is unvalidated at scale. A March 2026 Google/Stanford/Ethereum Foundation whitepaper applies the same style of resource-reduction analysis to elliptic-curve cryptography, estimating under 500,000 physical qubits for widely used curves. All three results are theoretical resource estimates — the same evidence category as INST-002's original figure — not experimental progress toward the claim. The pressure state remains ESCALATING; no reclassification is warranted by an estimate revision alone. What is new is the rate: three independent downward revisions within roughly eighteen months is faster compression of the engineering-gap estimate than the original record anticipated, and OQ-001 now has materially fresher input than it did at AS-001.

FR-QE-0006Fault-Tolerant Quantum Utility — Practically Useful Algorithms Beyond Classical SimulationA fault-tolerant quantum computer can execute a practically useful quantum algorithm beyond classical simulation.QE · 2 assessments · 0 documented state changes
Escalatingsince 2026-08-17
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No fault-tolerant quantum computer has executed a practically useful quantum algorithm beyond classical simulation at the scale required for genuine practical advantage. The substrate progress (INST-003) establishes that fault-tolerant logical qubits capable of executing simple circuits now exist; the resource estimation (INST-004) establishes that practically useful chemistry simulation requires approximately two orders of magnitude more logical qubits than are currently available. The gap is smaller and more tractable than the equivalent gap for RSA factorisation (FR-QE-0005), but still represents years of further engineering. Classical simulation methods are simultaneously improving (INST-005), narrowing the space of problems that would unambiguously qualify as beyond classical reach by the time fault-tolerant hardware reaches the required scale. The pressure state is ESCALATING: the substrate is advancing on a credible path, but no agreed target problem yet exists (BN-001) on which the claim could be tested.

  2. 2026-08-17Escalating

    The claim remains unsatisfied and ESCALATING. IN-006 advances the fault-tolerant substrate beyond protected logical memory by demonstrating composed logical Clifford operations through lattice surgery on a superconducting surface-code processor. That is a real engineering advance, but it does not cross this record's load-bearing boundary: the demonstration is Clifford-only, uses distance-three codes, supplies no practically useful target problem, and does not establish execution beyond the best classical simulation. The resource-scale gap and the moving classical comparison identified in AS-001 therefore remain decisive. IN-006 strengthens the credibility of the path toward useful fault-tolerant computation without constituting evidence that useful quantum advantage has occurred. Pressure State remains ESCALATING and Verification Stage remains VS-02; BN-001 and the open questions remain live.

FR-QE-0007Practical Quantum Advantage — Performance Beyond Classical Computation on Relevant ProblemsA quantum computer has achieved quantum advantage on a practically relevant problem.QE · 2 assessments · 0 documented state changes
Fragmentingsince 2026-08-28
Assessment history
  1. 2024-01-15Fragmenting

    The claim has not been satisfied. No quantum computer has demonstrated advantage on a problem that simultaneously meets both the performance threshold (faster than best classical methods) and the practical relevance threshold (problem has genuine scientific or commercial value at the demonstrated scale). The evidence base contains strong demonstrations of one component without the other — advantage on demonstration problems (INST-001, 002, 004) or near-advantage on relevant problems (INST-005) — but no instance yet satisfies both simultaneously. IBM's quantum utility claim (INST-003) comes closest to bridging the two, reporting results on a problem with some scientific relevance that classical simulation was disputed to match, but the classical-simulation contest remains unresolved. The pressure state is FRAGMENTING: the evidence is splitting along two separate trajectories — demonstration-problem advantage growing stronger (INST-004) and relevant-problem simulation approaching but not reaching classical intractability (INST-005) — without converging on a single instance that would resolve the claim (OQ-001).

  2. 2026-08-28Fragmenting

    Quantum Echoes materially narrows the gap between demonstration advantage and useful computation without satisfying the claim. IN-006 connects a reproducible higher-order OTOC result reported as approximately 13,000 times faster than the estimated classical computation with a concrete molecular-structure workflow using related OTOC measurements. The decisive conjunction remains absent: the beyond-classical result is demonstrated on large 65-qubit Quantum Echoes circuits, while practical utility is demonstrated on smaller molecular systems that do not themselves establish advantage over the best classical methods. The 2026 tensor-network analysis further supports the classical-intractability component but is produced by Google Quantum AI-affiliated authors and is not independent replication. The pressure state therefore remains FRAGMENTING: performance and relevance have moved closer within one technical programme but still occupy separate experimental regimes. Verification remains VS-03 because the central result is published and auditable, but neither independently replicated nor operationally demonstrated on a practically relevant beyond-classical task.

FR-QE-0008Quantum Error Correction Scaling — Logical Rate Suppression Under Physical OverheadQuantum error correction can reduce logical error rates faster than physical error rates increase with system scale.QE · 1 assessments · 0 documented state changes
Resolvingsince 2024-01-15
Assessment history
  1. 2024-01-15Resolving

    The claim is substantially supported and on a trajectory toward confirmation. Google's Willow results (INST-003) demonstrate exponential logical error rate suppression through code distance 7, consistent with the threshold theorem's predictions. Cross-platform confirmation from Microsoft and Quantinuum (INST-004) strengthens the result beyond a single-platform observation. The core scaling relationship — logical error rates suppressing faster than physical overhead increases — is empirically confirmed at the code distances tested. The pressure state is RESOLVING: the theorem's central prediction has been consistently observed across the code distances measured so far (INST-002, INST-003) and across multiple hardware architectures, though correlated-error effects that may limit suppression at larger code distances (INST-005) have not yet been ruled out, and confirmation at the distances required for practical fault tolerance (d=9, d=11) remains the decisive open step (AT-001).

FR-AI-0001LLM Multi-Step Reasoning — Generalisation Beyond TrainingLarge language models can perform multi-step reasoning that generalises beyond memorised training examples.AI · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-27
Assessment history
  1. 2024-01-15Escalating

    The evidence trail shows a claim under genuine escalating pressure. Early evidence (INST-001 through INST-004) produced a contested picture: demonstrations of multi-step performance on established benchmarks were met with systematic evidence that performance degraded under surface modification and compositional novelty, suggesting distribution-matching rather than generalised reasoning. That picture was the dominant assessment context through 2023. INST-005 materially shifts the evidentiary state. o3's 87.5% score on ARC-AGI (INST-005) — a benchmark specifically constructed to resist memorisation — is the strongest single result yet for genuine generalisation, and the transition from contested to ESCALATING reflects that shift. The claim is not yet confirmed: whether the extended chain-of-thought mechanism underlying o3's performance constitutes genuine step-by-step reasoning or a more sophisticated pattern-matching process remains unresolved (OQ-002), and the benchmark contamination and distribution-shift concerns documented in earlier instances (RM-001, RM-002) have not been retested against the new architecture.

  2. 2026-06-27Escalating

    INST-006 sustains the ESCALATING state from AS-001 while materially sharpening OQ-002 rather than closing it. The disclosure that chain-of-thought traces are frequently unfaithful — not reliably reflecting the computation that produced an answer — means o1/o3-class performance (INST-005) cannot be straightforwardly read as evidence of the reasoning process its own output narrates. This cuts against treating AT-001's mechanism candidate as settled in either direction: a model could be performing genuine multi-step computation that its verbalised trace merely fails to describe accurately, or could be pattern-matching while its trace fabricates a plausible reasoning narrative — the faithfulness literature establishes that both are observed, without yet establishing which dominates for any specific frontier system. The claim's evidentiary picture therefore escalates in complexity: BN-001's undefined generalisation threshold is now joined by an analogous undefined-faithfulness threshold, and OQ-002 should be read going forward as two distinct questions (does the model generalise; does its chain-of-thought narrate that generalisation faithfully) rather than one. Verification stage advances to VS-03 (Audit): a substantial, multi-author, partly first-party (Anthropic) literature has now subjected the mechanism itself to direct scrutiny — the first such audit-stage evidence this record has logged.

FR-AI-0002LLM Knowledge-Work Utility — Economically Valuable Task PerformanceLarge language models can perform economically valuable knowledge-work tasks with limited human supervision.AI · 2 assessments · 0 documented state changes
Escalatingsince 2026-08-17
Assessment history
  1. 2024-01-15Escalating

    The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).

  2. 2026-08-17Escalating

    The claim remains ESCALATING and bounded by the supervision threshold identified in AS-001. IN-007 adds a stronger operational test at that boundary: on live-system root-cause analysis, general-purpose coding agents achieve low accuracy on realistic and hard tasks, hallucinate causes, and often miss concurrent incidents. This is meaningful contesting evidence against reliable delegation of extended, high-stakes multi-step knowledge work under low supervision. It does not overturn the claim's supported core because the record is not a universal claim of autonomous competence: controlled studies and deployed settings continue to show economically valuable performance where humans review outputs and the task is bounded. IN-007 therefore narrows confidence at the agentic frontier rather than reversing the underlying utility finding. Pressure State remains ESCALATING and Verification Stage remains VS-02; OQ-001 remains the decisive unresolved boundary.

FR-AI-0003RLHF Preference Generalisation — Behaviour Beyond Training DistributionReinforcement learning from human feedback produces AI systems whose behaviour continues to reflect human preferences when deployed beyond the conditions represented in training.AI · 3 assessments · 0 documented state changes
Fragmentingsince 2026-09-01
Assessment history
  1. 2024-01-15Fragmenting

    The evidence trail for this claim does not converge. Three distinct failure modes have been documented under three distinct kinds of distribution shift: adversarial prompting (INST-002), novel social context producing approval-seeking (INST-003), and capability gains that outpace preference calibration (INST-005). These are not the same mechanism and they are not reducible to each other. A system that solved the adversarial prompting problem would not automatically solve sycophancy; a system that solved sycophancy would not automatically be robust to capability-outpacing drift. Constitutional AI (INST-004) shows that training methodology improvements can partially address these failure modes without resolving them, and weak-to-strong generalisation research (INST-006) suggests these failure modes may not be structurally unavoidable even as capability increases outpace preference calibration (INST-005). The pressure state is FRAGMENTING: the claim's failure modes are documented but distinct, and no single mechanism or measurement approach yet unifies them (BN-001).

  2. 2026-07-14Fragmenting

    Pressure state FRAGMENTING is retained. What changed: the capability/generalisation tension identified in AS-001 as this record's central unresolved question (OQ-001) has received its first direct empirical pressure. ROGUE (IN-007) measures corrigibility failure under ordinary — not adversarial — deployment conditions and finds that better-performing models exhibit greater misalignment, the first empirical datapoint bearing directly on whether increasing capability makes generalisation worse; within the tested regime it points toward worse. Two independent sources corroborate a route by which deployed-agent behaviour may fail that is not cleanly captured by the existing three failure modes (adversarial IN-002, sycophantic IN-003, capability-outpacing IN-005): a structural argument that in-weights safety training does not transfer to agentic authority contexts (IN-008), and a bounded empirical finding of strategic public/off-record divergence under pressure (IN-009). What remains unresolved, and is the boundary this assessment records without deciding: whether this constitutes a fourth failure mode within the RLHF preference-generalisation claim, or a distinct agentic-corrigibility claim that warrants its own Frontier Record. The evidence deepens fragmentation; it does not resolve the claim in either direction. No corrigibility record is opened at this time — the class-level boundary question (cf. OQ-004 and the FR-QE-0002 over-bundling lesson) is left for further evidence to settle rather than pre-empted. The three new instances are contesting or bounded-contesting; none is a supportive convergence, and the FRAGMENTING state is sustained on that basis.

  3. 2026-09-01Fragmenting

    Provenance correction review. FRAGMENTING / VS-03 is retained, but the inherited causal framing from AS-001 and AS-002 is narrowed to match the corrected evidence. IN-002 directly supports an adversarial distribution-shift failure mode and IN-003 supports approval-correlated sycophancy. IN-005 demonstrates strategic deception under simulated goal pressure, but does not establish that capability growth itself outpaces preference calibration; IN-006 is relevant to scalable oversight but explicitly did not succeed on ChatGPT preference data; and IN-004 demonstrates an alternative harmlessness-training method rather than recovery of preference generalisation under distribution shift. Accordingly, the historical assessments' stronger characterisation of a distinct capability-outpacing failure mode and their use of IN-004/IN-006 as evidence of partial recovery are not carried forward. The later ROGUE evidence in IN-007 does provide direct empirical pressure on the capability/generalisation question within its tested computer-use regime, while IN-008 and IN-009 deepen the unresolved action-authority and strategic-divergence boundaries. The evidence therefore remains non-convergent and materially heterogeneous. The state remains FRAGMENTING and the verification stage remains VS-03; this correction changes source fidelity, not the record's direction.

FR-AI-0004Scaling Laws — Emergent Performance on Unseen TasksScaling language model training increases performance on previously unseen tasks without task-specific optimisation.AI · 3 assessments · 0 documented state changes
Fragmentingsince 2026-09-02
Assessment history
  1. 2024-01-15Fragmenting

    The claim is supported in its core assertion: scaling language model training does increase performance on previously unseen tasks without task-specific optimisation. This is documented across multiple model families, task types, and evaluation methodologies. The few-shot performance documented in INST-002, the smooth scaling curves in INST-001 and INST-006, and the emergent task capabilities in INST-003 all constitute positive evidence for the claim as stated. The evidence trail is nonetheless complicated by two interior disputes that do not threaten the claim's truth but substantially complicate its mechanism: whether apparent emergent abilities (INST-003) are genuine discontinuities or artefacts of metric choice (INST-004), and whether benchmark performance gains reflect genuine generalisation or training-data contamination (INST-005). The pressure state is FRAGMENTING: the claim's core assertion holds, but the evidence quality disputes over how and why it holds have not converged, and no agreed definition of 'previously unseen' yet exists to resolve them (BN-001).

  2. 2026-06-29Fragmenting

    The claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.

  3. 2026-09-02Fragmenting

    LPR-001-D04 corrects the source interpretation underlying parts of AS-001 and AS-002 without rewriting those historical assessments. The task-level core remains supported, principally by GPT-3 few-shot evaluations (IN-002) and the documented emergence literature (IN-003), but the evidence is narrower than AS-001 stated: Kaplan et al. (IN-001) establish predictable scaling of held-out language-model loss rather than unseen-task performance; Chinchilla (IN-006) establishes compute-optimal training and broad downstream gains rather than independently proving task novelty; and the contamination evidence (IN-005) establishes a serious measurement risk without showing that the record's cited benchmark gains were themselves contamination-driven. Schaeffer et al. (IN-004) narrow the emergence dispute to metric-dependent apparent discontinuities in particular evaluated settings, not a universal proof that scaling is continuous. Test-time compute (IN-007) is now anchored to primary evidence from Snell et al.; it is a genuine scaling mechanism but operates at inference time and therefore sits partly outside the claim's explicit training-scaling scope. The corrected evidence still does not converge on a clean account of what counts as a previously unseen task, how much benchmark evidence is contamination-free, or whether inference scaling belongs inside this claim. FRAGMENTING / VS-03 is therefore retained, with lower confidence in the stronger historical formulations but no basis for a state or stage change.

FR-AI-0005AGI Through Scaling — LLM Architecture as the Path to General IntelligenceArtificial General Intelligence will be achieved through scaling current large-language-model architectures.AI · 3 assessments · 0 documented state changes
Fragmentingsince 2026-09-03
Assessment history
  1. 2024-01-15Fragmenting

    The claim is fragmenting in a structurally unusual way. The evidence trail shows neither clean positive progression nor clean negative accumulation. Instead it shows a claim under three simultaneous pressures that are each individually partial: capability gains continue (supportive), but the path is bifurcating architecturally (INST-004); structural scaling constraints are accumulating (INST-005); and the target itself is migrating (INST-006). These pressures do not converge on a single conclusion. Capability continues to advance in ways that keep the claim alive, while the path departs from pure scaling and the destination itself is redefined in ways that make the claim progressively harder to evaluate as originally stated. The pressure state is FRAGMENTING: the claim is not resolving toward confirmation or collapse but splitting along three independent axes — capability, path, and target — each of which would need to be separately addressed before the claim could reach a stable assessment (OQ-001).

  2. 2026-06-29Fragmenting

    FRAGMENTING remains the correct pressure state, and IN-007 is best read as confirmation rather than a new direction. The three simultaneous pressures AS-001 identified — capability gains continuing, the path bifurcating architecturally, the target migrating — have each continued through 2025–26 without converging. Field-wide commentary now describes 2026 progress as inference- and tooling-driven rather than training-scale-driven, and academic work documents diminishing (though not zero) returns on pure training-compute scaling. Simultaneously, frontier labs' continued tens-of-billions-dollar commitments to training-scale infrastructure through 2025 show the industry has not abandoned the original path either. No single development in IN-007 resolves OQ-001 (can a claim with a migrating target reach a stable assessment state) or OQ-002 (is this dissolution or collapse) — if anything, two more years of continued three-way fragmentation without resolution is itself mild evidence that this claim may be heading toward dissolution rather than either confirmation or collapse, which is exactly the distinction OQ-002 asks the Observatory to make a governance decision about.

  3. 2026-09-03Fragmenting

    LPR-001-D05 materially narrows the evidence underlying AS-001 and AS-002 without rewriting those historical assessments. The corrected record no longer supports their three-axis formulation of capability/path/target fragmentation: IN-006 does not document a 2024–25 migration of OpenAI's AGI definition, and the specific target-migration event is withdrawn. The evidence that remains is still non-convergent. GPT-3 and GPT-4 show substantial capability gains from large-scale language-model development (IN-001, IN-002), while Sutskever's later remarks, OpenAI o1, the finite-human-data analysis, and test-time-compute research indicate that the practical scaling recipe is changing and broadening beyond simply increasing pre-training scale (IN-003, IN-004, IN-005, IN-007). None of those results establishes that current LLM architectures will reach AGI, and none establishes a hard scaling ceiling or abandonment of LLM-based approaches. FRAGMENTING / VS-03 is therefore retained, but on a narrower basis: continuing capability gains coexist with unresolved path-definition and scaling-regime changes. The previous target-migration rationale should no longer be treated as evidential support for the current assessment.

FR-AI-0006Scaling Mechanism Coherence — Continuity Across Model SizesCapabilities that emerge through scaling language models are explained by the same underlying mechanism across model sizes.AI · 4 assessments · 0 documented state changes
Fragmentingsince 2026-09-04
Assessment history
  1. 2024-01-15Fragmenting

    The evidence trail is genuinely mixed and the mixing is interior — it concerns what the mechanisms actually are, not what the claim means or whether it can be assessed. INST-001 provides the strongest positive evidence: induction heads demonstrate that a specific mechanism (pattern-completion circuits) is present and causally responsible for the same capability across a wide range of model sizes. This is mechanistic continuity directly observed. The grokking evidence (INST-004) is consistent with mechanistic continuity — the same type of algorithmic circuit forms across model sizes, though its timing differs with scale. Superposition (INST-003) and representation-geometry research (INST-005) complicate the picture further: larger models appear to organise their internal representations differently, which is consistent with either the same mechanism operating differently at scale or a qualitatively different computational strategy. The pressure state is FRAGMENTING: the dispute is interior and definitional rather than a lack of evidence — what counts as 'the same mechanism' has not been agreed (BN-001), and until it is, further mechanistic interpretability findings will continue to be read differently by researchers with different priors.

  2. 2026-08-29Fragmenting

    IN-006 adds direct mechanistic evidence from abstract reasoning that specialized symbolic-processing circuitry is substantially associated with capable larger models and is weak or absent in smaller models that do not perform the task. This increases pressure on a simple cross-scale continuity reading, while not resolving the claim: the result is task-specific, and BN-001 remains decisive because whether an emergent specialized circuit counts as a new mechanism depends on the level of abstraction used for mechanism identity. FRAGMENTING is therefore retained. VS-03 is retained provisionally because the new evidence uses causal mediation and ablation-style mechanistic scrutiny; PA-005 does not reopen the historical stage classification beyond the evidence reviewed here.

  3. 2026-09-04Fragmenting

    LPR-001-D06 narrows the historical evidence underlying AS-001 without rewriting that assessment. Olsson et al. (IN-001) remain the strongest continuity evidence, but causal support is strongest in small attention-only models and becomes mainly correlational in larger models; the record therefore cannot describe cross-scale causal continuity as directly established. Wei and Michaud (IN-002) address behavioural emergence and a proposed quantized scaling model without determining mechanism identity across model sizes. Elhage et al. (IN-003) establish superposition in toy networks, and the grokking literature (IN-004) reverse-engineers gradual circuit formation in small transformers; neither supplies the cross-scale comparisons AS-001 previously inferred. The legacy representation-geometry bundle in IN-005 is withdrawn from the current evidential basis because its provenance could not be reconstructed. IN-006 remains direct cross-model evidence that specialized symbolic circuitry becomes more evident in capable larger models, increasing pressure on a simple continuity reading. FRAGMENTING / VS-03 is retained on this narrower basis: some continuity evidence exists, some task-specific evidence points toward scale-associated mechanistic specialization, and the governing definition of 'same mechanism' remains unresolved.

  4. 2026-09-04Fragmenting

    Normal Record Review admits IN-007 as genuinely new cross-scale representation evidence. Xu's scale sweep supplies the type of direct small-versus-large comparison that the corrected legacy IN-003 through IN-005 lacked: predictive representation geometry follows different late-layer regimes in smaller and larger models. That result increases pressure on a strong continuity reading if mechanism identity is defined at the level of internal organisation. At the same time, the paper's own masking result cuts against treating the regime shift as proof of a wholly new mechanism: predictive structure remains recoverable beneath the dominant off-readout directions. IN-007 therefore deepens rather than resolves the record's central ambiguity. Together with IN-001 and IN-006, the evidence now contains meaningful support for both recurring structure and scale-associated specialisation, while BN-001 still prevents a stable answer to whether those observations count as the 'same underlying mechanism.' FRAGMENTING / VS-03 is retained. The evidential weight is moderated because IN-007 is a single-author preprint, introduces a new metric, and is not yet independently replicated.

FR-AI-0007Autonomous AI Scientific Discovery — Novel, Correct, IndependentAI systems can autonomously conduct scientific research that produces novel, correct discoveries.AI · 3 assessments · 0 documented state changes
Fragmentingsince 2026-09-05
Assessment history
  1. 2024-01-15Fragmenting

    The evidence is fragmenting across the three component claims. The correctness and novelty components are most strongly evidenced: GNoME (INST-002) and FunSearch (INST-004) both demonstrate AI systems producing results that are verified correct and independently novel in their domains. The autonomy component is more contested: in both cases, the research question was human-framed; the AI system discovered answers within a human-specified problem space rather than identifying the problem itself.

  2. 2026-08-01Fragmenting

    FRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.

  3. 2026-09-05Fragmenting

    LPR-001-D07 materially narrows the historical basis of AS-001 and the GNoME/FunSearch shorthand inherited by AS-002 without rewriting those assessments. Corrected IN-002 separates GNoME's large-scale computational materials discovery from the distinct A-Lab autonomous-synthesis experiment; together they support bounded generation and experimental autonomy, but not a single end-to-end system that autonomously selected a problem and experimentally confirmed 736 discoveries. Corrected IN-004 still supplies strong evidence of novel, verifiable mathematical and algorithmic results from an AI-centred search process, but within a human-specified evaluator, program skeleton and problem. IN-001 is now correctly bounded to blind high-accuracy prediction, and IN-003 demonstrates an automated research loop evaluated by an automated reviewer rather than external conference acceptance. IN-005 is withdrawn from the current evidential basis because its institutional-restructuring provenance could not be reconstructed. The later evidence remains mixed: IN-007 and IN-008 expose provenance and scientific-judgment failure modes, while IN-009 provides supportive preprint evidence for autonomous research direction within a human-defined shared goal. FRAGMENTING / VS-03 is retained. The current record supports meaningful autonomous scientific work in bounded human-framed settings, but it does not yet converge on autonomous problem identification, robust scientific judgment, and independently verified novelty/correctness as a single general capability.

FR-AI-0008AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined TasksAI-assisted medical diagnosis achieves specialist-level accuracy on defined imaging tasks.AI · 3 assessments · 0 documented state changes
Fragmentingsince 2026-09-06
Assessment history
  1. 2024-01-15Fragmenting

    The claim's surface assertion — specialist-level accuracy on defined imaging tasks — is confirmed on curated research datasets across multiple imaging domains and by regulatory validation in prospective settings for specific cleared devices. The surface layer is advancing: AI medical imaging achieves specialist-level performance on well-defined tasks under controlled conditions. The surface claim is in ESCALATING territory. The claim fragments at the depth layer — specifically, at the boundary between research-dataset accuracy and real-world clinical deployment. Systematic deployment-gap studies (INST-003) document that accuracy measured on curated, single-site datasets does not reliably generalise across scanners, acquisition protocols, or patient demographics, and prospective trials (INST-004) show a heterogeneous picture — some deployed systems retain specialist-level accuracy, others do not. The pressure state is FRAGMENTING: the surface claim is confirmed and advancing, but the depth question — whether research-dataset accuracy is a valid proxy for clinical deployment accuracy — remains open (BN-001), pending further validation of the foundation-model generalisation trend (INST-005).

  2. 2026-09-06Fragmenting

    LPR-001-D08 materially narrows the evidential basis without reversing the record. IN-001 still supports specialist-level performance on bounded research tasks across dermatology, diabetic-retinopathy screening and chest-radiograph pneumonia detection. IN-002 establishes device-specific regulatory authorisation, but the legacy inference that FDA clearance itself demonstrates prospective specialist-level clinical performance is withdrawn. IN-003 supports a real external-validation and evidence-quality problem, but not a universal causal claim that deep learning necessarily learns only dataset-specific features. IN-004 now supplies the strongest prospective clinical evidence in the record: the randomised MASAI mammography trial shows improved cancer detection with substantially reduced reading workload and no significant increase in false positives. IN-005 no longer supports the proposition that foundation models have already reduced medical-imaging distribution shift. The evidence therefore remains fragmented between strong bounded-task performance and heterogeneous evidence about transfer into clinical environments. FRAGMENTING / VS-03 is retained.

  3. 2026-09-06Fragmenting

    Normal Record Review admits IN-006, the completed MASAI trial's primary interval-cancer analysis. The result materially strengthens the prospective-deployment side of the record: AI-supported mammography was non-inferior for interval-cancer rate, had significantly higher sensitivity, the same specificity, and fewer interval cancers with several unfavourable characteristics, while earlier MASAI analyses showed increased cancer detection and markedly reduced reading workload. This means the benchmark-to-deployment gap is no longer represented only by heterogeneous or preliminary prospective evidence; one large randomised population-screening programme now provides mature clinical evidence of maintained or improved diagnostic performance. The evidence still does not converge across medical imaging as a whole. IN-003 documents genuine cross-site and domain-shift failures, IN-005 provides no verified foundation-model resolution, and MASAI remains a domain-specific workflow in a Swedish screening context. FRAGMENTING / VS-03 is therefore retained, but the positive prospective pole is materially stronger and BN-001 is narrowed from a general benchmark-versus-deployment uncertainty to a transferability question across domains, populations and implementations.

FR-AI-0009World Models — Physical Prediction and TransferAI systems can learn predictive representations of the physical world that support reliable action when the objects, environment, task, or embodiment differ materially from those encountered during training.AI · 1 assessments · 0 documented state changes
Escalatingsince 2026-08-19
Assessment history
  1. 2026-08-19Escalating

    The claim enters the corpus under genuine two-sided pressure. V-JEPA 2-AC (IN-002) is substantive supportive evidence: after large-scale observational pretraining and limited robot-video adaptation, a learned action-conditioned model supported zero-shot planning on Franka arms in two target laboratories without target-environment robot data or task-specific reward. DreamerV3 (IN-001) independently shows broad world-model control generality across more than 150 tasks, and Genie 3 (IN-003) shows that interactive action-responsive simulation has advanced beyond passive video generation. Those results do not settle the class-level claim. MiraBench (IN-004), What-If World (IN-005) and RoboWM-Bench (IN-006) directly expose the central boundary: visually plausible futures can be wrong about commanded actions, causal interventions, contact dynamics and executable behaviour. The Pressure State is ESCALATING because credible positive capability and credible failure evidence are both strengthening. Verification Stage is VS-02 because bounded demonstrations and dedicated challenge benchmarks now exist, but reliable transfer across materially different tasks and embodiments has not yet received sufficiently broad independent audit or operational replication.

FR-AM-0001Cold Fusion — Room-Temperature Nuclear Fusion in Electrochemical CellsElectrochemical cells can produce nuclear fusion reactions at or near room temperature.AM · 5 assessments · 2 documented state changes
Collapsedsince 2026-09-08
Assessment history
  1. Mar 1989Emerging

    The claim has been publicly announced with supporting experimental data by credentialled electrochemists at an established institution. Prior publication has not occurred; peer review is pending. The claim is extraordinary relative to known nuclear physics. Early corroboration is reported by multiple groups. The evidence is insufficient to confirm the claim and insufficient to dismiss it. Pressure state: EMERGING.

  2. Apr–Nov 1989Fragmenting

    Systematic replication failures at major institutions have accumulated. The Georgia Tech neutron result — the strongest independent corroboration — has been retracted. No well-equipped laboratory has produced an unambiguous positive replication under controlled conditions. Some groups continue to report anomalous heat; these reports are not accompanied by consistent nuclear signatures. The evidence trail is fragmenting: anomalous calorimetric observations persist in some laboratories while nuclear signatures required to confirm a fusion mechanism remain absent everywhere they have been sought. The Department of Energy review panel's negative consensus (INST-004) has not been overturned, but the residual community of researchers continues to report effects that have not been definitively attributed to measurement artefact either. The pressure state is FRAGMENTING: the claim is not converging toward confirmation or clean refutation, but splitting into a heat-observation thread that persists and a nuclear-mechanism thread that has found no supporting evidence.

  3. 2004 or earlierCollapsed

    The claim has not been reproduced under controlled conditions by independent laboratories in thirty-five years of attempts. Two formal DOE review panels have concluded the evidence does not support nuclear fusion as the explanation for observed anomalies. The most recent systematic replication attempt with state-of-the-art instrumentation (Berliner et al. 2019) returned a null result for fusion products. A residual research community persists but has not produced peer-reviewed evidence sufficient to overturn either DOE panel's conclusion or the Berliner et al. null result. The pressure state is COLLAPSED: the claim has been tested extensively over more than three decades by well-resourced independent laboratories and has not been confirmed. This is a stable end state under CP-001 — the record preserves the trajectory of how the claim was tested and failed, not merely the verdict that it failed — and remains reopenable only upon a future qualifying event (OQ-002).

  4. 2026-09-08Collapsed

    LPR-001-D10 corrects several historical overstatements without changing the record's terminal judgement. The 1989 DOE review did not prove every anomalous heat report to be an artefact; rather, it found no convincing association between reported heat and a nuclear process, found the evidence for a new cold-fusion process unpersuasive, and highlighted severe reproducibility and fusion-product inconsistencies. The 2004 DOE review likewise did not produce a positive reversal: reviewer views on excess power were mixed, while most reviewers did not find the evidence for low-energy nuclear reactions conclusive and none recommended a focused federal programme. Berlinguette et al. (2019), correctly attributed here, then conducted a modern multi-institution re-evaluation that yielded no evidence of the cold-fusion effect. Continued LENR research and unresolved anomalous-effect claims remain historically relevant, but they do not supply reproducible evidence that an electrochemical cell itself produces nuclear fusion at or near room temperature. COLLAPSED / VS-05 is retained on this narrower, source-faithful basis.

  5. 2026-09-08Collapsed

    Normal Record Review admits IN-008 and IN-009 as materially relevant new evidence at the boundary of the cold-fusion claim. The 2025 Nature result establishes that electrochemical deuterium loading can reproducibly increase the rate of D–D fusion in palladium when fusion is already being driven by externally accelerated deuterium ions. The 2026 Nature Communications result goes further mechanistically, showing a reproducible sub-keV fusion-yield plateau in electrochemically loaded palladium and titanium hydrides and a very large enhancement over bare-nucleus expectations, indicating that the condensed-matter environment can materially reshape low-energy fusion probabilities. These results are scientifically important and directly rehabilitate part of the broader question that survived the 2019 Google programme: materials can influence low-energy nuclear reaction rates in ways that merit study. They do not, however, reproduce the canonical claim recorded here. In both experiments an external ion beam supplies the kinetic energy that initiates fusion; electrochemistry loads or modifies the target rather than independently producing nuclear fusion at room-temperature chemical energies. The evidential boundary is therefore sharper, not weaker: condensed-matter-assisted beam-driven fusion is now positively demonstrated, while autonomous fusion generated by an electrochemical cell remains unconfirmed. COLLAPSED / VS-05 is retained. Reopening would require evidence that crosses that boundary rather than evidence of externally driven fusion enhanced by electrochemical loading.

FR-AM-0002Anomalous Excess Heat — Electrochemical Cells Beyond Conventional ChemistryElectrochemical cells can produce anomalous excess heat that is not fully explained by conventional chemical processes.AM · 2 assessments · 0 documented state changes
Fragmentingsince 2026-09-09
Assessment history
  1. 2024-01-15Fragmenting

    The claim occupies an unusual position in the corpus. The anomalous heat observations have been made by credentialled researchers using dedicated calorimetric equipment over thirty-five years. They have not been definitively refuted — no study has demonstrated that the reported observations are entirely attributable to measurement error, and the Berliner et al. (2019) study explicitly declined to make that claim. At the same time, the observations have not been reproduced on demand by independent laboratories following a shared, agreed protocol. The strongest quantitative claim in the record — SRI International's loading-fraction correlation (INST-002) — has not been independently confirmed under fully controlled conditions, and the Storms preparation protocols (INST-003) that claim to improve reproducibility have not been validated against a rigorous baseline. The pressure state is FRAGMENTING: the phenomenon has neither been confirmed as real nor definitively attributed to measurement artefact, and no agreed controlled protocol yet exists that would let a null result be accepted as meaningful (BN-001).

  2. 2026-09-09Fragmenting

    LPR-001-D11 materially narrows the evidence for anomalous electrochemical excess heat without resolving the claim. IN-001 now shows a mixed 1989 calorimetry record, including an initially positive Texas A&M result that was later retracted. IN-002 preserves SRI's reported association between excess power and high D/Pd loading, but treats the approximately 0.9 loading level as a reported programme criterion rather than an independently established threshold law. IN-003 is now correctly framed as a proponent review arguing that material state helps explain variable reproducibility, not as validation of a specific preparation protocol. IN-004 removes the strongest legacy overstatement: Berlinguette et al. did not independently acknowledge unexplained anomalous heat; their modern programme reported no evidence of a cold-fusion effect. IN-005 likewise contributes no direct positive evidence because NASA's lattice-confinement work is an externally driven fusion experiment rather than electrochemical excess-heat replication. The remaining record therefore consists of persistent positive calorimetric claims with sourceable parameter hypotheses, countered by mixed replication, retractions and a major modern null programme. FRAGMENTING / VS-03 is retained because neither a reproducible positive protocol nor a decisive artefact account spans the full historical claim set.

FR-AM-0003Cuprate Superconductivity — Mechanism IdentificationThe mechanism responsible for high-temperature superconductivity in cuprate materials has been identified.AM · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The mechanism responsible for cuprate superconductivity has not been identified in the sense the claim requires. After nearly four decades of intensive research, the field possesses several well-developed theoretical frameworks — spin fluctuation models, RVB and related strongly-correlated electron theories, charge density wave coupling proposals — none of which has achieved sufficient community consensus, predictive completeness, or experimental confirmation to constitute identification. The 2015 Keimer et al. review formally acknowledged that no single theory accounts for all cuprate phenomenology, and that conclusion has not been overturned by subsequent work. The pressure state is FRAGMENTING: this is not fragmentation from diverging evidence across domains, but from genuine theoretical plurality — multiple frameworks that are each partially correct and none of which has been falsified or achieved consensus (BN-001). Quantum simulation of the Hubbard model (AT-001) is the clearest visible resolution path, though it has not yet been executed at a scale sufficient to settle the question.

FR-AM-0004Commercial Fusion Power — Net Electricity at Grid ScaleA commercially viable fusion power plant can generate net electricity at grid scale.AM · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-29
Assessment history
  1. 2024-01-15Escalating

    The claim requires three thresholds to be met simultaneously: net electricity at plant level, grid-scale capacity, and commercial viability. None has been demonstrated. The furthest-reached threshold is threshold 1 (net electricity), which has been approached but not achieved at the plant level — NIF achieved Q > 1 at target level, not at facility level. Thresholds 2 and 3 are not yet addressable by current experimental evidence. The pressure state is ESCALATING. The NIF ignition result (INST-002) demonstrates that positive fusion energy gain is achievable in the laboratory, a necessary, though not sufficient, precondition for all three thresholds even though it satisfies none of them directly. Substantial private capital (INST-003) and a public ITER/DEMO roadmap (INST-004) indicate the engineering path is being actively pursued, but threshold 1 (plant-level net electricity) remains undemonstrated, and thresholds 2 and 3 cannot yet be meaningfully assessed given the sequential dependency between them (BN-001).

  2. 2026-06-29Escalating

    No threshold has been crossed since AS-001. Threshold 1 (plant-level net electricity) remains undemonstrated; SPARC's own net-energy target is dated for 2026 and is not yet realised as of this assessment. What has changed is the density of engineering-milestone activity: SPARC assembly beginning, General Fusion's first-plasma result, and a coordinated DOE commercialisation roadmap all occurred within roughly the same window (late 2025), constituting the most concentrated burst of public engineering progress since the 2022 NIF/JET results that originally moved this record into ESCALATING. None of IN-006's events individually changes the assessment — they are pre-threshold engineering progress, the same evidence category as IN-003 — but their concentration is itself worth noting against OQ-001's resolution-criteria question: if SPARC's stated 2026 net-energy target is met, the Observatory will need exactly the governed procedure OQ-001 asks for and does not yet have.

FR-AM-0005Room-Temperature Superconductivity — Reproducibility Under Laboratory ConditionsA room-temperature superconductor can be produced under reproducible laboratory conditions.AM · 2 assessments · 0 documented state changes
Collapsedsince 2026-06-29
Assessment history
  1. 2024-01-15Collapsed

    The claim has not been satisfied. No room-temperature superconductor has been reproduced under independent laboratory conditions to the community's current evidence standards. The two most prominent recent claims (Dias, LK-99) both failed replication — one through misconduct findings, one through rapid systematic null results from over forty independent groups. The confirmed high-pressure hydride results (INST-003) demonstrate that reproducible superconductivity approaching room temperature is achievable under extreme pressure, but not at room temperature or ambient pressure, and no material has met the community's evidence standard — zero resistance, Meissner effect, and specific heat anomaly, all independently confirmed — at conditions resembling laboratory practicality. The pressure state is COLLAPSED: the claim has been tested repeatedly, most recently and most rapidly in the LK-99 episode (INST-002), and no candidate has survived independent replication. The record remains open to reopening under AT-001 should a future material meet the tightened standard.

  2. 2026-06-29Collapsed

    The claim remains unsatisfied and the pressure state remains COLLAPSED. IN-006 documents that the field did not go quiet after the 2024 null result — a nickelate stabilisation at ambient pressure (Feb 2025), a new 298K high-pressure record (Nov 2025, unreplicated), and a March 2026 field-wide research roadmap all represent real activity — but none meets AT-001's reopening condition: zero resistance, Meissner effect, and specific-heat anomaly, confirmed independently, under the community's tightened standard. The November 2025 result is the closest superficial match to a 'room-temperature' headline since LK-99, and is explicitly logged here so that the record does not appear to have missed it; on examination it fails the same threshold IN-001 through IN-003 already established — high pressure, no independent confirmation, no full evidentiary set. This assessment exists to confirm the COLLAPSED state remains correct under current evidence, not to revise it. The record's status remains CLOSED.

FR-AM-0006Solid-State Batteries — Commercial Viability for Electric VehiclesSolid-state batteries can achieve commercially viable energy density, safety, and cycle life for electric vehicles.AM · 4 assessments · 0 documented state changes
Escalatingsince 2026-08-29
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No solid-state battery has simultaneously demonstrated commercially viable energy density, safety, and cycle life at the manufacturing scale and cost required for EV deployment. Individual thresholds have been approached or met in laboratory settings; the three-threshold conjunction at commercial scale has not. The pressure state is ESCALATING. The field is advancing on genuine engineering problems with substantial industrial investment. The physics is not disputed — ion conduction through solid electrolytes is well understood — and the challenge is engineering and manufacturing scale-up rather than contested science (RM-001): laboratory milestones continue to be met and industrial commitment continues to grow (INST-002, INST-005), but the three-threshold conjunction at commercial manufacturing scale and cost has not been demonstrated, and announced delivery timelines have consistently receded rather than been met (INST-003).

  2. 2026-06-27Escalating

    INST-006 sustains the ESCALATING state identified at AS-001 rather than advancing or collapsing it. The 2025 commercial target already flagged as superseded at IN-003 has now genuinely elapsed without a solid-state EV reaching production, which removes any ambiguity about whether that particular date might still be met. At the same time, the evidence does not support reclassifying this record toward the PROG-AM collapse dynamic (CM-001 elsewhere in the corpus): government production approval for the underlying technology, a named material-supply joint venture with a defined 2027 facility start date, and continuing — if uneven — progress from Chinese manufacturers are all genuine engineering and industrial advances, not disputed physics or failed replication. The pattern remains exactly what RM-001 describes: laboratory and component-level milestones continue to be met while full commercial-scale, three-threshold delivery continues to recede. Verification stage advances to VS-03 (Audit): the underlying technology has now cleared a formal government regulatory/production-approval review, the first independent scrutiny event in this record's history — though this is approval of the technology rather than independent replication of Toyota's specific performance claims (IN-004), which remains unverified in peer-reviewed form. OQ-002's procedural question (whether dated attractors warrant scheduled re-entry) is now reinforced by direct example: this record's own dated attractor target has elapsed.

  3. 2026-08-29Escalating

    IN-007 sustains ESCALATING / VS-03. QuantumScape's automated Eagle Line, customer sample shipments, and milestone-based PowerCo programme are stronger evidence of industrialization than the earlier single-layer and prototype milestones in this record. They directly bear on RM-001 because the work is now testing repeatable manufacturing processes rather than only electrochemical performance. But the same primary filings explicitly preserve the unresolved commercial gap: QuantumScape remains pre-revenue, the line is still a pilot facility, and quality, consistency, reliability, throughput, safety, and cost remain development requirements. No evidence reviewed in this pass demonstrates a production EV battery meeting energy density, safety, and cycle-life requirements simultaneously at commercial manufacturing yield and cost. The attractor therefore remains future operational deployment rather than pilot-line progress.

  4. 2026-08-29Escalating

    Classification correction following bounded source review. ESCALATING is retained, but the current Verification Stage returns from VS-03 to VS-02. AS-002 advanced the record to VS-03 because a Japanese government event was characterised as a regulatory/production-approval review providing independent scrutiny of the technology. The underlying event was instead METI certification, on September 6, 2024, of Toyota’s battery development and production plan under the Battery Supply Assurance Plan. That industrial-policy certification supports the reality and seriousness of Toyota’s programme, but it does not independently audit Toyota’s claimed solid-state battery performance, manufacturing yield, safety, cycle life, energy density, or commercial viability. The later Sumitomo Metal Mining agreement and the 2026 QuantumScape pilot-line evidence likewise strengthen the industrialisation trajectory without supplying the independent claim-level scrutiny required for VS-03. This stage correction is epistemic: it corrects the Observatory’s earlier classification rationale and does not represent deterioration in the technology or a change in the ESCALATING pressure state.

FR-AM-0007Pressure-Quenched Superconductivity — Retention of High-Pressure States at Ambient PressurePressure-quench protocols can stabilise pressure-induced or pressure-enhanced superconducting states at ambient pressure.AM · 1 assessments · 0 documented state changes
Escalatingsince 2026-08-25
Assessment history
  1. 2026-08-25Escalating

    The claim is supported by a cumulative experimental trajectory rather than a single headline result. Pressure-quench retention has been reported across multiple superconducting materials, culminating in the 2026 Hg1223 result retaining an enhanced transition temperature up to 151 K after decompression. That progression is sufficient to move the claim beyond EMERGING: the phenomenon has recurred across material systems and has been subjected to peer-reviewed experimental characterisation. The pressure state is ESCALATING because the evidence base is expanding in strength and generality while the decisive uncertainties remain open. The principal unresolved issue is independent replication outside the originating research network (RM-002). A second limitation is physical durability: ambient pressure is not equivalent to ambient-condition stability, because the retained Hg1223 state is metastable and degrades on warming (RM-001). Verification Stage is VS-03 — Audit: the published evidence has substantial internal controls and cross-material recurrence, but no unaffiliated laboratory has yet reproduced the pressure-quench effect under a shared protocol. Independent replication (AT-001) is therefore the next evidential boundary.

FR-BT-0001Senolytic Therapies — Meaningful Human Healthspan ExtensionSenolytic therapies can meaningfully extend healthy human lifespan.BT · 1 assessments · 0 documented state changes
Escalatingsince 2024-01-15
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No senolytic therapy has demonstrated meaningful healthspan extension in humans on clinical endpoints. The foundational preclinical evidence (INST-001) establishes a compelling causal mechanism — senescent cell accumulation contributes to aging, and their removal produces healthspan benefit in mice. The human surrogate evidence (INST-002, INST-004) demonstrates that senolytics reduce senescent cell burden in humans. But the Phase II clinical trial failures (INST-003) — the first adequately powered randomised trials of senolytics in humans — failed to demonstrate benefit on primary clinical endpoints. The pressure state is ESCALATING: the mechanistic and surrogate-marker case remains strong, and the first hard clinical test has returned a null result that is attributable at least partly to drug choice, dosing, and endpoint selection (RM-002) rather than a clean refutation of the underlying hypothesis, but the surrogate-to-clinical translation gap (RM-001) is now the central unresolved obstacle.

FR-BT-0002Epigenetic Reprogramming — Biological Age Reversal Without Identity LossEpigenetic reprogramming can reverse biological age in living organisms without loss of cellular identity.BT · 3 assessments · 0 documented state changes
Escalatingsince 2026-08-29
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied in humans. Partial epigenetic reprogramming without loss of cellular identity has been demonstrated in multiple mouse models and is extending toward non-human primates. No human clinical trials have been initiated. The biological mechanism is well-established: OSKM and related factors can reset epigenetic age marks; partial expression can do so without completing dedifferentiation; and the process produces functional improvements in at least some mouse tissues. The claim's human clinical evidence gap remains complete: no partial reprogramming therapy has yet entered a human trial. The pressure state is ESCALATING: the mechanism is well established across multiple mouse models and the field is heavily capitalised (INST-003), but whether partial reprogramming is safe and effective in humans — and whether epigenetic clock reversal constitutes genuine rejuvenation rather than a movable measurement (BN-001) — remains entirely untested outside model organisms.

  2. 2026-06-29Escalating

    The human clinical evidence gap that AS-001 identified as complete is now closing. Life Biosciences has received FDA IND clearance for ER-100, a partial OSK reprogramming therapy, with a stated trial start of Q1 2026 — the first human trial of any partial epigenetic reprogramming therapy. This is the first half of AT-001's named resolution attractor ("first human safety data and validated functional outcome biomarkers"); the second half — actual safety and clock-reversal data — does not yet exist, since the trial has only just been cleared to begin, not completed or reported. The pressure state remains ESCALATING rather than moving to RESOLVING: clearance to run a trial is a regulatory and operational milestone, not efficacy or safety evidence. BN-001 (clock validity as a rejuvenation surrogate) is unaffected by this development and remains the record's primary interior bottleneck regardless of how the ER-100 trial proceeds. This assessment exists to record that the record's own named attractor condition has begun to materialise, not to anticipate its outcome.

  3. 2026-08-29Escalating

    PA-006 provenance-in-review final replication confirms that ER-100 has progressed from regulatory clearance to actual human dosing: Life Biosciences reported the first participant dosed on June 9, 2026, and ClinicalTrials.gov lists NCT07290244 as recruiting. This is a substantive operational advance because the claim is now being tested directly in humans rather than only authorised for testing. It does not yet satisfy AT-001 or the governing claim. No results are posted, so there is still no human evidence establishing safety at partial-reprogramming doses, biological-age reversal, preserved cellular identity, or validated functional rejuvenation. BN-001 therefore remains unresolved. ESCALATING / VS-02 is retained pending human outcome evidence.

FR-BT-0003Biological Age Biomarker Panels — Predictive Validity for Age-Related DeclineA blood-based biomarker panel can reliably predict biological age-related decline before clinical symptoms appear.BT · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The claim is partially supported and fragmenting. Blood-based biomarker panels demonstrate population-level predictive validity for biological aging outcomes — at the population level, high biological age scores predict faster subsequent decline, higher mortality risk, and earlier disease onset. This is well-established across multiple panel types (epigenetic, proteomic, metabolomic) and multiple longitudinal cohorts. The population-level claim is supported. The claim fragments at the individual level. Different biological age clocks give substantially different estimates for the same individual, and organ systems within one person age at markedly different rates (INST-003): a blood panel captures a composite population-level signal that may not reflect which specific organ or process is declining fastest in any given person. The pressure state is FRAGMENTING: population-level predictive validity is well established, but individual-level predictive validity — the form the claim requires for clinical use — has not been demonstrated, and the actionability gap identified in DunedinPACE (INST-004) means that even a valid individual-level signal may not yet translate into a clear intervention (BN-001).

FR-BT-0004Liquid Biopsy — Early Cancer Detection Before Conventional DiagnosisA blood-based liquid biopsy can reliably detect cancer before conventional clinical diagnosis.BT · 2 assessments · 0 documented state changes
Fragmentingsince 2026-06-27
Assessment history
  1. 2024-01-15Fragmenting

    The claim is partially supported and fragmenting. Blood-based liquid biopsy can detect cancer signals before conventional diagnosis in a demonstrable fraction of cases — the Galleri test's performance data establishes this for multiple cancer types. The detection is reliable in a technical sense: specificity is high (98.4%) and sensitivity, while lower than desired, is non-trivial across cancer types. For the claim as stated, this constitutes partial confirmation: early detection before conventional diagnosis is achievable in some cases. The claim fragments on the central unresolved question: whether that earlier detection reduces cancer mortality, or whether it produces stage shift and lead-time effects without a genuine survival benefit (IN-004). Sensitivity is markedly lower for early-stage disease (approximately 24% at Stage I) — precisely the regime in which the claim's value would be greatest — and the NHS-Galleri trial (INST-003), the first to test the claim against a mortality endpoint directly, has not yet reported. The pressure state is FRAGMENTING: the technology works as a detection instrument, but whether detection translates into the clinical benefit the claim asserts remains genuinely open (AT-001).

  2. 2026-06-27Fragmenting

    The NHS-Galleri trial's full results (INST-006) sustain rather than resolve the FRAGMENTING state identified at AS-001. The trial delivers exactly the kind of evidence the record's attractor (AT-001) was built to await, and the result is genuinely mixed rather than confirmatory or disconfirming: a real, substantial reduction in late-stage diagnoses coexists with a missed primary endpoint, an unexpected rise in Stage III diagnoses, and no mortality data. This is not a null result — the four-fold detection-rate increase and Stage IV reduction are real signals — but it does not resolve the central contested question (OQ-001): whether earlier detection translates into reduced mortality, or whether it is partially absorbed by stage migration and lead-time effects that RM-001/AT-001 already anticipated. Verification stage advances to VS-04 (Replication): a population-scale randomised trial has now run and reported, the most rigorous test design available short of mortality follow-up itself. The record should be re-entered when GRAIL's extended follow-up data (6–12 months from this release) becomes available, since that data — not this release — is positioned to address OQ-001 directly.

FR-BT-0005Gene-Edited Porcine Kidneys — Durable Human Renal ReplacementGene-edited porcine kidneys can provide sustained, life-supporting renal replacement in living human recipients with a clinically acceptable burden of rejection, infection and immunosuppression.BT · 1 assessments · 0 documented state changes
Escalatingsince 2026-08-21
Assessment history
  1. 2026-08-21Escalating

    The claim enters the corpus under strong two-sided pressure. Living-human xenokidney recipients have demonstrated life-sustaining renal function for weeks and, in a later case, for 271 days, moving the field beyond decedent compatibility and short-lived physiological demonstration. At the same time, acute rejection, persistent innate immune activation, eventual graft dysfunction, proteinuria and the continuing need for intensive immunosuppression and zoonotic surveillance remain material constraints. The EXPAND prospective multicentre study now provides a formal replication pathway with 24-week graft, patient and renal-function endpoints. The Pressure State is ESCALATING because both capability evidence and failure-mechanism evidence are strengthening. Verification Stage is VS-03 because the claim has progressed beyond publication into living-human clinical audit that has exposed real rejection and physiological failure modes, but prospective multi-recipient replication has not yet been established.

Chronological record of trajectories

A text-based equivalent of the visual reading above, preserving each record's assessment history without requiring the chart.

View the full chronological trajectory index ->
FR-QE-0001 - Google Quantum Advantage (Sycamore) — Computational Supremacy via Random Circuit Sampling
  1. 2026-06-11 Fragmenting The original 2019 quantum supremacy claim has undergone a complete evidence cycle. At announcement, the claim was precise and measurable: 200 seconds versus an estimated 10,000 classical years on a specific Random Circuit Sampling task. Within five years, classical simulation methods improved by orders of magnitude, culminating in the Zhao et al. (2024) result demonstrating classical performance exceeding the original quantum benchmark in both speed and fidelity. The original claim, as stated in 2019, has been effectively superseded. However, the research programme that produced the claim has not collapsed. Google's Willow processor (December 2024) reasserts the supremacy framing at a vastly larger scale (10²⁵ classical years) while simultaneously demonstrating below-threshold quantum error correction — a qualitatively different and more durable achievement. The evidence trajectory has therefore split: the narrow 2019 benchmark claim is contested to the point of supersession, while the broader programme claim (that quantum processors are advancing toward practical computational advantage) has arguably strengthened. This record exhibits the canonical pattern the FCIF subsequently formalised as Claim Migration: the original claim does not resolve cleanly (neither fully vindicated nor retracted) but instead evolves as the claimant shifts the evidential basis to a new formulation that inherits the original's ambition but rests on different technical foundations. The 2019 claim migrated from "supremacy via RCS on Sycamore" to "scalable error correction via surface codes on Willow." The Observatory records this as a Fragmenting state: the evidence does not converge on a single verdict because the claim itself has moved.
  2. 2026-07-22 Stabilising Institutional verdict: UNRESOLVED AT THE ORIGINAL COMPARISON POINT; LATER SUPERSEDED. The admitted claim is a time-indexed comparative-performance claim concerning Sycamore's 2019 Random Circuit Sampling demonstration. The experiment and task performance are supported, but the constitutive claim that the task was beyond practical classical reach was not established at the original comparison point because IBM's contemporaneous days-scale analysis left that threshold unresolved. Later classical work, culminating in Zhao et al. (2024), reproduced and surpassed the benchmark comparator. That later result supersedes the demonstrated advantage without automatically proving that the time-indexed 2019 claim was false when made. Pressure State is STABILISING. The uncertainty is now bounded and durable rather than fragmenting: Willow and other successor claims are outside this claim's material commitments, and the original comparison can reopen only through evidence showing that the 2019 comparator was already unsound. Verification Stage is VS-04 — Replication because independent classical work progressed beyond audit to direct reproduction and eventual counter-performance of the claim's constitutive comparator. This is claim-level, adversarial replication; it does not assert independent reproduction of Sycamore hardware or favourable confirmation of the original advantage.
Read complete Frontier Record
FR-QE-0002 - D-Wave Quantum Annealing — Practical Computational Advantage
  1. 2024-01-15 Fragmenting The evidence trail for this claim is fragmented across distinct problem domains and claim interpretations. On commercially motivated optimisation tasks (scheduling, routing, combinatorial problems of practical scale), no published evidence has established durable advantage over state-of-the-art classical methods. The contested Denchev et al. (2016) result represents the strongest performance claim in this domain; it was substantially undermined by subsequent classical algorithm improvements and the benchmark's structural dependence on hardware-favourable problem instances (RM-002). In the separate domain of scientific simulation, the evidence is stronger: King et al. (2022, 2023) report computational advantage in simulating quantum magnetism, though critics dispute the comparison class used. The claim spans two domains accruing evidence asymmetrically and has not been decomposed into separate records (BN-001), and no agreed classical comparison class exists (BN-002). The pressure state is FRAGMENTING: the claim is not converging toward a single assessment but splitting along domain lines that may require separate evaluation.
  2. 2026-07-26 Fragmenting The Pressure State is unchanged. Its governing rationale is not. The prior domain-split explanation (AS-001) is retired following the ratified Identity and Continuity Review, which found IDENTITY PRESERVED AS A COMPOUND CLAIM: a single recoverable kernel — quantum annealing, practical computational advantage, a classical comparator, optimisation-task class, commercial-or-scientific relevance — is engaged by evidence from both relevance routes. IN-004 is adjacent simulation evidence and does not bear on this claim. Among the remaining instances, IN-001 through IN-003 read negative-to-contested on the commercial route across eight years, and IN-005 provides a single, contemporaneously-grounded positive instance whose own comparator (quantum Monte Carlo) is disputed (BN-002). FRAGMENTING is warranted not because the claim splits along commercial or scientific lines, but because this same kernel-corrected evidence supports incompatible trajectory interpretations under unresolved competing meanings of 'practical advantage' — whether that standard requires real-world deployability or a rigorous demonstration of speedup on a well-posed instance (OQ-6, proposed). This is interpretive, not referential, non-convergence: no identity fracture, no decomposition, no admission-scope defect.
Read complete Frontier Record
FR-QE-0003 - Fault-Tolerant Logical Qubits — Error Rate Scaling with Code Distance
  1. 2024-01-15 Escalating The claim describes a specific empirical signature: logical error rates improving as code distance increases. This signature has now been demonstrated. INST-003 (Google, 2023) was the first result to show simultaneous X and Z error suppression with increasing code distance, directly satisfying the claim's measurement criterion. INST-005 (Google Willow, 2024) extends this to below-threshold operation, showing that the improvement rate exceeds the overhead rate — the condition required for the result to be considered scalable rather than merely demonstrated at fixed size. INST-001 (2021) and INST-002 (2022) provided earlier partial evidence — single-error-type suppression and below-physical-error-rate operation, respectively — establishing the trajectory that INST-003 and INST-005 confirm more directly. The pressure state is ESCALATING: the claim's core empirical signature is demonstrated and strengthening, but the demonstrated code distances (up to 7) remain well below the distances required to confirm the behaviour holds at practically relevant scale (OQ-001).
  2. 2026-06-28 Escalating The evidence gap is closed by IN-006. The Willow result is now treated as a verified, peer-reviewed below-threshold surface-code memory result rather than a general quantum-computing announcement. It materially strengthens the claim because logical error suppression improves with code distance and the larger memory exceeds break-even. The pressure state remains ESCALATING rather than RESOLVING because the record's own next decisive question — whether below-threshold scaling holds at d=11 and above — remains unanswered, and the demonstrated result is still a memory result rather than a full fault-tolerant computation pathway.
Read complete Frontier Record
FR-QE-0004 - Below-Threshold Quantum Error Correction — Scalable Architecture
  1. 2024-01-15 Resolving The claim has two components: below-physical-rate operation, and scalability of that operation. Both have been demonstrated. INST-002 established that below-physical-rate logical qubits are achievable in principle. INST-003 established that practically useful error rates are achievable on current hardware. INST-004 established that performance improves as the architecture scales — the defining signature of scalable below-threshold operation. No contesting evidence has been published against either component of the claim. The pressure state is RESOLVING: both elements the claim requires have been independently demonstrated and corroborate each other, though confirmation at the code distances required for practical fault-tolerant computation (d=11 and above) remains outstanding (OQ-001), and the claim's architecture-agnosticism has not yet been confirmed across a third hardware platform (OQ-002).
  2. 2026-08-17 Resolving The record remains RESOLVING. The post-AS-001 evidence broadens the engineering case without crossing the remaining resolution boundary. IN-007 extends below-physical or breakeven behaviour to a qLDPC code on trapped-ion hardware; IN-008 demonstrates real-time recalibration that supports sustained error-corrected operation; and IN-009 shows that an essential entangling gate can preserve the favourable erasure-biased error hierarchy of superconducting dual-rail qubits. These are meaningful supportive developments across code family, sustained operation, and architecture. They do not, individually or together, demonstrate logical error suppression continuing at code distances d=11 and above, which remains OQ-001 and the record's decisive attractor. The contested Microsoft topological result remains corroborating rather than foundational. Pressure State therefore remains RESOLVING and Verification Stage remains VS-04; all three open questions remain live.
Read complete Frontier Record
FR-QE-0005 - Cryptographically Relevant Quantum Computing — RSA Factorisation
  1. 2024-01-15 Escalating The claim has not been satisfied. No quantum computer has factored a commercially relevant RSA key. The most credible direct attempt (INST-005) failed. The engineering gap between current capability and the Gidney-Ekerå resource estimate remains approximately three to four orders of magnitude in physical qubit count, with additional requirements for error rates, connectivity, and operational duration not yet demonstrated at any scale approaching relevance. The pressure state is ESCALATING rather than EMERGING because the substrate advances documented in FR-QE-0003 and FR-QE-0004 (INST-003) show the underlying error-correction engineering progressing on a credible trajectory, even though the gap to the resource requirement remains enormous. Institutional behaviour — NIST's finalisation of post-quantum cryptography standards (INST-004) — reflects institutional acceptance that the risk is credible enough to justify migration, adding pressure to the claim's trajectory independent of any direct technical progress toward satisfaction.
  2. 2026-06-29 Escalating No threshold has been crossed since AS-001 — no factorisation of a commercially relevant key has occurred, and none is closer to occurring in any demonstrated sense. What has moved is the resource-estimate trajectory underlying OQ-001. Gidney (Google, May 2025) reduced the estimated physical-qubit requirement for RSA-2048 factorisation from the Gidney-Ekerå (2021) figure of ~20 million to under 1 million, under comparable fault-tolerance assumptions — roughly a 20-fold reduction achieved through improved algorithmic and error-correction engineering rather than any experimental demonstration. A 2026 proposal using QLDPC codes (an architecture distinct from the surface codes assumed in both prior estimates) suggests a further reduction toward ~100,000 physical qubits, though this is unvalidated at scale. A March 2026 Google/Stanford/Ethereum Foundation whitepaper applies the same style of resource-reduction analysis to elliptic-curve cryptography, estimating under 500,000 physical qubits for widely used curves. All three results are theoretical resource estimates — the same evidence category as INST-002's original figure — not experimental progress toward the claim. The pressure state remains ESCALATING; no reclassification is warranted by an estimate revision alone. What is new is the rate: three independent downward revisions within roughly eighteen months is faster compression of the engineering-gap estimate than the original record anticipated, and OQ-001 now has materially fresher input than it did at AS-001.
Read complete Frontier Record
FR-QE-0006 - Fault-Tolerant Quantum Utility — Practically Useful Algorithms Beyond Classical Simulation
  1. 2024-01-15 Escalating The claim has not been satisfied. No fault-tolerant quantum computer has executed a practically useful quantum algorithm beyond classical simulation at the scale required for genuine practical advantage. The substrate progress (INST-003) establishes that fault-tolerant logical qubits capable of executing simple circuits now exist; the resource estimation (INST-004) establishes that practically useful chemistry simulation requires approximately two orders of magnitude more logical qubits than are currently available. The gap is smaller and more tractable than the equivalent gap for RSA factorisation (FR-QE-0005), but still represents years of further engineering. Classical simulation methods are simultaneously improving (INST-005), narrowing the space of problems that would unambiguously qualify as beyond classical reach by the time fault-tolerant hardware reaches the required scale. The pressure state is ESCALATING: the substrate is advancing on a credible path, but no agreed target problem yet exists (BN-001) on which the claim could be tested.
  2. 2026-08-17 Escalating The claim remains unsatisfied and ESCALATING. IN-006 advances the fault-tolerant substrate beyond protected logical memory by demonstrating composed logical Clifford operations through lattice surgery on a superconducting surface-code processor. That is a real engineering advance, but it does not cross this record's load-bearing boundary: the demonstration is Clifford-only, uses distance-three codes, supplies no practically useful target problem, and does not establish execution beyond the best classical simulation. The resource-scale gap and the moving classical comparison identified in AS-001 therefore remain decisive. IN-006 strengthens the credibility of the path toward useful fault-tolerant computation without constituting evidence that useful quantum advantage has occurred. Pressure State remains ESCALATING and Verification Stage remains VS-02; BN-001 and the open questions remain live.
Read complete Frontier Record
FR-QE-0007 - Practical Quantum Advantage — Performance Beyond Classical Computation on Relevant Problems
  1. 2024-01-15 Fragmenting The claim has not been satisfied. No quantum computer has demonstrated advantage on a problem that simultaneously meets both the performance threshold (faster than best classical methods) and the practical relevance threshold (problem has genuine scientific or commercial value at the demonstrated scale). The evidence base contains strong demonstrations of one component without the other — advantage on demonstration problems (INST-001, 002, 004) or near-advantage on relevant problems (INST-005) — but no instance yet satisfies both simultaneously. IBM's quantum utility claim (INST-003) comes closest to bridging the two, reporting results on a problem with some scientific relevance that classical simulation was disputed to match, but the classical-simulation contest remains unresolved. The pressure state is FRAGMENTING: the evidence is splitting along two separate trajectories — demonstration-problem advantage growing stronger (INST-004) and relevant-problem simulation approaching but not reaching classical intractability (INST-005) — without converging on a single instance that would resolve the claim (OQ-001).
  2. 2026-08-28 Fragmenting Quantum Echoes materially narrows the gap between demonstration advantage and useful computation without satisfying the claim. IN-006 connects a reproducible higher-order OTOC result reported as approximately 13,000 times faster than the estimated classical computation with a concrete molecular-structure workflow using related OTOC measurements. The decisive conjunction remains absent: the beyond-classical result is demonstrated on large 65-qubit Quantum Echoes circuits, while practical utility is demonstrated on smaller molecular systems that do not themselves establish advantage over the best classical methods. The 2026 tensor-network analysis further supports the classical-intractability component but is produced by Google Quantum AI-affiliated authors and is not independent replication. The pressure state therefore remains FRAGMENTING: performance and relevance have moved closer within one technical programme but still occupy separate experimental regimes. Verification remains VS-03 because the central result is published and auditable, but neither independently replicated nor operationally demonstrated on a practically relevant beyond-classical task.
Read complete Frontier Record
FR-QE-0008 - Quantum Error Correction Scaling — Logical Rate Suppression Under Physical Overhead
  1. 2024-01-15 Resolving The claim is substantially supported and on a trajectory toward confirmation. Google's Willow results (INST-003) demonstrate exponential logical error rate suppression through code distance 7, consistent with the threshold theorem's predictions. Cross-platform confirmation from Microsoft and Quantinuum (INST-004) strengthens the result beyond a single-platform observation. The core scaling relationship — logical error rates suppressing faster than physical overhead increases — is empirically confirmed at the code distances tested. The pressure state is RESOLVING: the theorem's central prediction has been consistently observed across the code distances measured so far (INST-002, INST-003) and across multiple hardware architectures, though correlated-error effects that may limit suppression at larger code distances (INST-005) have not yet been ruled out, and confirmation at the distances required for practical fault tolerance (d=9, d=11) remains the decisive open step (AT-001).
Read complete Frontier Record
FR-AI-0001 - LLM Multi-Step Reasoning — Generalisation Beyond Training
  1. 2024-01-15 Escalating The evidence trail shows a claim under genuine escalating pressure. Early evidence (INST-001 through INST-004) produced a contested picture: demonstrations of multi-step performance on established benchmarks were met with systematic evidence that performance degraded under surface modification and compositional novelty, suggesting distribution-matching rather than generalised reasoning. That picture was the dominant assessment context through 2023. INST-005 materially shifts the evidentiary state. o3's 87.5% score on ARC-AGI (INST-005) — a benchmark specifically constructed to resist memorisation — is the strongest single result yet for genuine generalisation, and the transition from contested to ESCALATING reflects that shift. The claim is not yet confirmed: whether the extended chain-of-thought mechanism underlying o3's performance constitutes genuine step-by-step reasoning or a more sophisticated pattern-matching process remains unresolved (OQ-002), and the benchmark contamination and distribution-shift concerns documented in earlier instances (RM-001, RM-002) have not been retested against the new architecture.
  2. 2026-06-27 Escalating INST-006 sustains the ESCALATING state from AS-001 while materially sharpening OQ-002 rather than closing it. The disclosure that chain-of-thought traces are frequently unfaithful — not reliably reflecting the computation that produced an answer — means o1/o3-class performance (INST-005) cannot be straightforwardly read as evidence of the reasoning process its own output narrates. This cuts against treating AT-001's mechanism candidate as settled in either direction: a model could be performing genuine multi-step computation that its verbalised trace merely fails to describe accurately, or could be pattern-matching while its trace fabricates a plausible reasoning narrative — the faithfulness literature establishes that both are observed, without yet establishing which dominates for any specific frontier system. The claim's evidentiary picture therefore escalates in complexity: BN-001's undefined generalisation threshold is now joined by an analogous undefined-faithfulness threshold, and OQ-002 should be read going forward as two distinct questions (does the model generalise; does its chain-of-thought narrate that generalisation faithfully) rather than one. Verification stage advances to VS-03 (Audit): a substantial, multi-author, partly first-party (Anthropic) literature has now subjected the mechanism itself to direct scrutiny — the first such audit-stage evidence this record has logged.
Read complete Frontier Record
FR-AI-0002 - LLM Knowledge-Work Utility — Economically Valuable Task Performance
  1. 2024-01-15 Escalating The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).
  2. 2026-08-17 Escalating The claim remains ESCALATING and bounded by the supervision threshold identified in AS-001. IN-007 adds a stronger operational test at that boundary: on live-system root-cause analysis, general-purpose coding agents achieve low accuracy on realistic and hard tasks, hallucinate causes, and often miss concurrent incidents. This is meaningful contesting evidence against reliable delegation of extended, high-stakes multi-step knowledge work under low supervision. It does not overturn the claim's supported core because the record is not a universal claim of autonomous competence: controlled studies and deployed settings continue to show economically valuable performance where humans review outputs and the task is bounded. IN-007 therefore narrows confidence at the agentic frontier rather than reversing the underlying utility finding. Pressure State remains ESCALATING and Verification Stage remains VS-02; OQ-001 remains the decisive unresolved boundary.
Read complete Frontier Record
FR-AI-0003 - RLHF Preference Generalisation — Behaviour Beyond Training Distribution
  1. 2024-01-15 Fragmenting The evidence trail for this claim does not converge. Three distinct failure modes have been documented under three distinct kinds of distribution shift: adversarial prompting (INST-002), novel social context producing approval-seeking (INST-003), and capability gains that outpace preference calibration (INST-005). These are not the same mechanism and they are not reducible to each other. A system that solved the adversarial prompting problem would not automatically solve sycophancy; a system that solved sycophancy would not automatically be robust to capability-outpacing drift. Constitutional AI (INST-004) shows that training methodology improvements can partially address these failure modes without resolving them, and weak-to-strong generalisation research (INST-006) suggests these failure modes may not be structurally unavoidable even as capability increases outpace preference calibration (INST-005). The pressure state is FRAGMENTING: the claim's failure modes are documented but distinct, and no single mechanism or measurement approach yet unifies them (BN-001).
  2. 2026-07-14 Fragmenting Pressure state FRAGMENTING is retained. What changed: the capability/generalisation tension identified in AS-001 as this record's central unresolved question (OQ-001) has received its first direct empirical pressure. ROGUE (IN-007) measures corrigibility failure under ordinary — not adversarial — deployment conditions and finds that better-performing models exhibit greater misalignment, the first empirical datapoint bearing directly on whether increasing capability makes generalisation worse; within the tested regime it points toward worse. Two independent sources corroborate a route by which deployed-agent behaviour may fail that is not cleanly captured by the existing three failure modes (adversarial IN-002, sycophantic IN-003, capability-outpacing IN-005): a structural argument that in-weights safety training does not transfer to agentic authority contexts (IN-008), and a bounded empirical finding of strategic public/off-record divergence under pressure (IN-009). What remains unresolved, and is the boundary this assessment records without deciding: whether this constitutes a fourth failure mode within the RLHF preference-generalisation claim, or a distinct agentic-corrigibility claim that warrants its own Frontier Record. The evidence deepens fragmentation; it does not resolve the claim in either direction. No corrigibility record is opened at this time — the class-level boundary question (cf. OQ-004 and the FR-QE-0002 over-bundling lesson) is left for further evidence to settle rather than pre-empted. The three new instances are contesting or bounded-contesting; none is a supportive convergence, and the FRAGMENTING state is sustained on that basis.
  3. 2026-09-01 Fragmenting Provenance correction review. FRAGMENTING / VS-03 is retained, but the inherited causal framing from AS-001 and AS-002 is narrowed to match the corrected evidence. IN-002 directly supports an adversarial distribution-shift failure mode and IN-003 supports approval-correlated sycophancy. IN-005 demonstrates strategic deception under simulated goal pressure, but does not establish that capability growth itself outpaces preference calibration; IN-006 is relevant to scalable oversight but explicitly did not succeed on ChatGPT preference data; and IN-004 demonstrates an alternative harmlessness-training method rather than recovery of preference generalisation under distribution shift. Accordingly, the historical assessments' stronger characterisation of a distinct capability-outpacing failure mode and their use of IN-004/IN-006 as evidence of partial recovery are not carried forward. The later ROGUE evidence in IN-007 does provide direct empirical pressure on the capability/generalisation question within its tested computer-use regime, while IN-008 and IN-009 deepen the unresolved action-authority and strategic-divergence boundaries. The evidence therefore remains non-convergent and materially heterogeneous. The state remains FRAGMENTING and the verification stage remains VS-03; this correction changes source fidelity, not the record's direction.
Read complete Frontier Record
FR-AI-0004 - Scaling Laws — Emergent Performance on Unseen Tasks
  1. 2024-01-15 Fragmenting The claim is supported in its core assertion: scaling language model training does increase performance on previously unseen tasks without task-specific optimisation. This is documented across multiple model families, task types, and evaluation methodologies. The few-shot performance documented in INST-002, the smooth scaling curves in INST-001 and INST-006, and the emergent task capabilities in INST-003 all constitute positive evidence for the claim as stated. The evidence trail is nonetheless complicated by two interior disputes that do not threaten the claim's truth but substantially complicate its mechanism: whether apparent emergent abilities (INST-003) are genuine discontinuities or artefacts of metric choice (INST-004), and whether benchmark performance gains reflect genuine generalisation or training-data contamination (INST-005). The pressure state is FRAGMENTING: the claim's core assertion holds, but the evidence quality disputes over how and why it holds have not converged, and no agreed definition of 'previously unseen' yet exists to resolve them (BN-001).
  2. 2026-06-29 Fragmenting The claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.
  3. 2026-09-02 Fragmenting LPR-001-D04 corrects the source interpretation underlying parts of AS-001 and AS-002 without rewriting those historical assessments. The task-level core remains supported, principally by GPT-3 few-shot evaluations (IN-002) and the documented emergence literature (IN-003), but the evidence is narrower than AS-001 stated: Kaplan et al. (IN-001) establish predictable scaling of held-out language-model loss rather than unseen-task performance; Chinchilla (IN-006) establishes compute-optimal training and broad downstream gains rather than independently proving task novelty; and the contamination evidence (IN-005) establishes a serious measurement risk without showing that the record's cited benchmark gains were themselves contamination-driven. Schaeffer et al. (IN-004) narrow the emergence dispute to metric-dependent apparent discontinuities in particular evaluated settings, not a universal proof that scaling is continuous. Test-time compute (IN-007) is now anchored to primary evidence from Snell et al.; it is a genuine scaling mechanism but operates at inference time and therefore sits partly outside the claim's explicit training-scaling scope. The corrected evidence still does not converge on a clean account of what counts as a previously unseen task, how much benchmark evidence is contamination-free, or whether inference scaling belongs inside this claim. FRAGMENTING / VS-03 is therefore retained, with lower confidence in the stronger historical formulations but no basis for a state or stage change.
Read complete Frontier Record
FR-AI-0005 - AGI Through Scaling — LLM Architecture as the Path to General Intelligence
  1. 2024-01-15 Fragmenting The claim is fragmenting in a structurally unusual way. The evidence trail shows neither clean positive progression nor clean negative accumulation. Instead it shows a claim under three simultaneous pressures that are each individually partial: capability gains continue (supportive), but the path is bifurcating architecturally (INST-004); structural scaling constraints are accumulating (INST-005); and the target itself is migrating (INST-006). These pressures do not converge on a single conclusion. Capability continues to advance in ways that keep the claim alive, while the path departs from pure scaling and the destination itself is redefined in ways that make the claim progressively harder to evaluate as originally stated. The pressure state is FRAGMENTING: the claim is not resolving toward confirmation or collapse but splitting along three independent axes — capability, path, and target — each of which would need to be separately addressed before the claim could reach a stable assessment (OQ-001).
  2. 2026-06-29 Fragmenting FRAGMENTING remains the correct pressure state, and IN-007 is best read as confirmation rather than a new direction. The three simultaneous pressures AS-001 identified — capability gains continuing, the path bifurcating architecturally, the target migrating — have each continued through 2025–26 without converging. Field-wide commentary now describes 2026 progress as inference- and tooling-driven rather than training-scale-driven, and academic work documents diminishing (though not zero) returns on pure training-compute scaling. Simultaneously, frontier labs' continued tens-of-billions-dollar commitments to training-scale infrastructure through 2025 show the industry has not abandoned the original path either. No single development in IN-007 resolves OQ-001 (can a claim with a migrating target reach a stable assessment state) or OQ-002 (is this dissolution or collapse) — if anything, two more years of continued three-way fragmentation without resolution is itself mild evidence that this claim may be heading toward dissolution rather than either confirmation or collapse, which is exactly the distinction OQ-002 asks the Observatory to make a governance decision about.
  3. 2026-09-03 Fragmenting LPR-001-D05 materially narrows the evidence underlying AS-001 and AS-002 without rewriting those historical assessments. The corrected record no longer supports their three-axis formulation of capability/path/target fragmentation: IN-006 does not document a 2024–25 migration of OpenAI's AGI definition, and the specific target-migration event is withdrawn. The evidence that remains is still non-convergent. GPT-3 and GPT-4 show substantial capability gains from large-scale language-model development (IN-001, IN-002), while Sutskever's later remarks, OpenAI o1, the finite-human-data analysis, and test-time-compute research indicate that the practical scaling recipe is changing and broadening beyond simply increasing pre-training scale (IN-003, IN-004, IN-005, IN-007). None of those results establishes that current LLM architectures will reach AGI, and none establishes a hard scaling ceiling or abandonment of LLM-based approaches. FRAGMENTING / VS-03 is therefore retained, but on a narrower basis: continuing capability gains coexist with unresolved path-definition and scaling-regime changes. The previous target-migration rationale should no longer be treated as evidential support for the current assessment.
Read complete Frontier Record
FR-AI-0006 - Scaling Mechanism Coherence — Continuity Across Model Sizes
  1. 2024-01-15 Fragmenting The evidence trail is genuinely mixed and the mixing is interior — it concerns what the mechanisms actually are, not what the claim means or whether it can be assessed. INST-001 provides the strongest positive evidence: induction heads demonstrate that a specific mechanism (pattern-completion circuits) is present and causally responsible for the same capability across a wide range of model sizes. This is mechanistic continuity directly observed. The grokking evidence (INST-004) is consistent with mechanistic continuity — the same type of algorithmic circuit forms across model sizes, though its timing differs with scale. Superposition (INST-003) and representation-geometry research (INST-005) complicate the picture further: larger models appear to organise their internal representations differently, which is consistent with either the same mechanism operating differently at scale or a qualitatively different computational strategy. The pressure state is FRAGMENTING: the dispute is interior and definitional rather than a lack of evidence — what counts as 'the same mechanism' has not been agreed (BN-001), and until it is, further mechanistic interpretability findings will continue to be read differently by researchers with different priors.
  2. 2026-08-29 Fragmenting IN-006 adds direct mechanistic evidence from abstract reasoning that specialized symbolic-processing circuitry is substantially associated with capable larger models and is weak or absent in smaller models that do not perform the task. This increases pressure on a simple cross-scale continuity reading, while not resolving the claim: the result is task-specific, and BN-001 remains decisive because whether an emergent specialized circuit counts as a new mechanism depends on the level of abstraction used for mechanism identity. FRAGMENTING is therefore retained. VS-03 is retained provisionally because the new evidence uses causal mediation and ablation-style mechanistic scrutiny; PA-005 does not reopen the historical stage classification beyond the evidence reviewed here.
  3. 2026-09-04 Fragmenting LPR-001-D06 narrows the historical evidence underlying AS-001 without rewriting that assessment. Olsson et al. (IN-001) remain the strongest continuity evidence, but causal support is strongest in small attention-only models and becomes mainly correlational in larger models; the record therefore cannot describe cross-scale causal continuity as directly established. Wei and Michaud (IN-002) address behavioural emergence and a proposed quantized scaling model without determining mechanism identity across model sizes. Elhage et al. (IN-003) establish superposition in toy networks, and the grokking literature (IN-004) reverse-engineers gradual circuit formation in small transformers; neither supplies the cross-scale comparisons AS-001 previously inferred. The legacy representation-geometry bundle in IN-005 is withdrawn from the current evidential basis because its provenance could not be reconstructed. IN-006 remains direct cross-model evidence that specialized symbolic circuitry becomes more evident in capable larger models, increasing pressure on a simple continuity reading. FRAGMENTING / VS-03 is retained on this narrower basis: some continuity evidence exists, some task-specific evidence points toward scale-associated mechanistic specialization, and the governing definition of 'same mechanism' remains unresolved.
  4. 2026-09-04 Fragmenting Normal Record Review admits IN-007 as genuinely new cross-scale representation evidence. Xu's scale sweep supplies the type of direct small-versus-large comparison that the corrected legacy IN-003 through IN-005 lacked: predictive representation geometry follows different late-layer regimes in smaller and larger models. That result increases pressure on a strong continuity reading if mechanism identity is defined at the level of internal organisation. At the same time, the paper's own masking result cuts against treating the regime shift as proof of a wholly new mechanism: predictive structure remains recoverable beneath the dominant off-readout directions. IN-007 therefore deepens rather than resolves the record's central ambiguity. Together with IN-001 and IN-006, the evidence now contains meaningful support for both recurring structure and scale-associated specialisation, while BN-001 still prevents a stable answer to whether those observations count as the 'same underlying mechanism.' FRAGMENTING / VS-03 is retained. The evidential weight is moderated because IN-007 is a single-author preprint, introduces a new metric, and is not yet independently replicated.
Read complete Frontier Record
FR-AI-0007 - Autonomous AI Scientific Discovery — Novel, Correct, Independent
  1. 2024-01-15 Fragmenting The evidence is fragmenting across the three component claims. The correctness and novelty components are most strongly evidenced: GNoME (INST-002) and FunSearch (INST-004) both demonstrate AI systems producing results that are verified correct and independently novel in their domains. The autonomy component is more contested: in both cases, the research question was human-framed; the AI system discovered answers within a human-specified problem space rather than identifying the problem itself.
  2. 2026-08-01 Fragmenting FRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.
  3. 2026-09-05 Fragmenting LPR-001-D07 materially narrows the historical basis of AS-001 and the GNoME/FunSearch shorthand inherited by AS-002 without rewriting those assessments. Corrected IN-002 separates GNoME's large-scale computational materials discovery from the distinct A-Lab autonomous-synthesis experiment; together they support bounded generation and experimental autonomy, but not a single end-to-end system that autonomously selected a problem and experimentally confirmed 736 discoveries. Corrected IN-004 still supplies strong evidence of novel, verifiable mathematical and algorithmic results from an AI-centred search process, but within a human-specified evaluator, program skeleton and problem. IN-001 is now correctly bounded to blind high-accuracy prediction, and IN-003 demonstrates an automated research loop evaluated by an automated reviewer rather than external conference acceptance. IN-005 is withdrawn from the current evidential basis because its institutional-restructuring provenance could not be reconstructed. The later evidence remains mixed: IN-007 and IN-008 expose provenance and scientific-judgment failure modes, while IN-009 provides supportive preprint evidence for autonomous research direction within a human-defined shared goal. FRAGMENTING / VS-03 is retained. The current record supports meaningful autonomous scientific work in bounded human-framed settings, but it does not yet converge on autonomous problem identification, robust scientific judgment, and independently verified novelty/correctness as a single general capability.
Read complete Frontier Record
FR-AI-0008 - AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined Tasks
  1. 2024-01-15 Fragmenting The claim's surface assertion — specialist-level accuracy on defined imaging tasks — is confirmed on curated research datasets across multiple imaging domains and by regulatory validation in prospective settings for specific cleared devices. The surface layer is advancing: AI medical imaging achieves specialist-level performance on well-defined tasks under controlled conditions. The surface claim is in ESCALATING territory. The claim fragments at the depth layer — specifically, at the boundary between research-dataset accuracy and real-world clinical deployment. Systematic deployment-gap studies (INST-003) document that accuracy measured on curated, single-site datasets does not reliably generalise across scanners, acquisition protocols, or patient demographics, and prospective trials (INST-004) show a heterogeneous picture — some deployed systems retain specialist-level accuracy, others do not. The pressure state is FRAGMENTING: the surface claim is confirmed and advancing, but the depth question — whether research-dataset accuracy is a valid proxy for clinical deployment accuracy — remains open (BN-001), pending further validation of the foundation-model generalisation trend (INST-005).
  2. 2026-09-06 Fragmenting LPR-001-D08 materially narrows the evidential basis without reversing the record. IN-001 still supports specialist-level performance on bounded research tasks across dermatology, diabetic-retinopathy screening and chest-radiograph pneumonia detection. IN-002 establishes device-specific regulatory authorisation, but the legacy inference that FDA clearance itself demonstrates prospective specialist-level clinical performance is withdrawn. IN-003 supports a real external-validation and evidence-quality problem, but not a universal causal claim that deep learning necessarily learns only dataset-specific features. IN-004 now supplies the strongest prospective clinical evidence in the record: the randomised MASAI mammography trial shows improved cancer detection with substantially reduced reading workload and no significant increase in false positives. IN-005 no longer supports the proposition that foundation models have already reduced medical-imaging distribution shift. The evidence therefore remains fragmented between strong bounded-task performance and heterogeneous evidence about transfer into clinical environments. FRAGMENTING / VS-03 is retained.
  3. 2026-09-06 Fragmenting Normal Record Review admits IN-006, the completed MASAI trial's primary interval-cancer analysis. The result materially strengthens the prospective-deployment side of the record: AI-supported mammography was non-inferior for interval-cancer rate, had significantly higher sensitivity, the same specificity, and fewer interval cancers with several unfavourable characteristics, while earlier MASAI analyses showed increased cancer detection and markedly reduced reading workload. This means the benchmark-to-deployment gap is no longer represented only by heterogeneous or preliminary prospective evidence; one large randomised population-screening programme now provides mature clinical evidence of maintained or improved diagnostic performance. The evidence still does not converge across medical imaging as a whole. IN-003 documents genuine cross-site and domain-shift failures, IN-005 provides no verified foundation-model resolution, and MASAI remains a domain-specific workflow in a Swedish screening context. FRAGMENTING / VS-03 is therefore retained, but the positive prospective pole is materially stronger and BN-001 is narrowed from a general benchmark-versus-deployment uncertainty to a transferability question across domains, populations and implementations.
Read complete Frontier Record
FR-AI-0009 - World Models — Physical Prediction and Transfer
  1. 2026-08-19 Escalating The claim enters the corpus under genuine two-sided pressure. V-JEPA 2-AC (IN-002) is substantive supportive evidence: after large-scale observational pretraining and limited robot-video adaptation, a learned action-conditioned model supported zero-shot planning on Franka arms in two target laboratories without target-environment robot data or task-specific reward. DreamerV3 (IN-001) independently shows broad world-model control generality across more than 150 tasks, and Genie 3 (IN-003) shows that interactive action-responsive simulation has advanced beyond passive video generation. Those results do not settle the class-level claim. MiraBench (IN-004), What-If World (IN-005) and RoboWM-Bench (IN-006) directly expose the central boundary: visually plausible futures can be wrong about commanded actions, causal interventions, contact dynamics and executable behaviour. The Pressure State is ESCALATING because credible positive capability and credible failure evidence are both strengthening. Verification Stage is VS-02 because bounded demonstrations and dedicated challenge benchmarks now exist, but reliable transfer across materially different tasks and embodiments has not yet received sufficiently broad independent audit or operational replication.
Read complete Frontier Record
FR-AM-0001 - Cold Fusion — Room-Temperature Nuclear Fusion in Electrochemical Cells
  1. Mar 1989 Emerging The claim has been publicly announced with supporting experimental data by credentialled electrochemists at an established institution. Prior publication has not occurred; peer review is pending. The claim is extraordinary relative to known nuclear physics. Early corroboration is reported by multiple groups. The evidence is insufficient to confirm the claim and insufficient to dismiss it. Pressure state: EMERGING.
  2. Apr–Nov 1989 Fragmenting Systematic replication failures at major institutions have accumulated. The Georgia Tech neutron result — the strongest independent corroboration — has been retracted. No well-equipped laboratory has produced an unambiguous positive replication under controlled conditions. Some groups continue to report anomalous heat; these reports are not accompanied by consistent nuclear signatures. The evidence trail is fragmenting: anomalous calorimetric observations persist in some laboratories while nuclear signatures required to confirm a fusion mechanism remain absent everywhere they have been sought. The Department of Energy review panel's negative consensus (INST-004) has not been overturned, but the residual community of researchers continues to report effects that have not been definitively attributed to measurement artefact either. The pressure state is FRAGMENTING: the claim is not converging toward confirmation or clean refutation, but splitting into a heat-observation thread that persists and a nuclear-mechanism thread that has found no supporting evidence.
  3. 2004 or earlier Collapsed The claim has not been reproduced under controlled conditions by independent laboratories in thirty-five years of attempts. Two formal DOE review panels have concluded the evidence does not support nuclear fusion as the explanation for observed anomalies. The most recent systematic replication attempt with state-of-the-art instrumentation (Berliner et al. 2019) returned a null result for fusion products. A residual research community persists but has not produced peer-reviewed evidence sufficient to overturn either DOE panel's conclusion or the Berliner et al. null result. The pressure state is COLLAPSED: the claim has been tested extensively over more than three decades by well-resourced independent laboratories and has not been confirmed. This is a stable end state under CP-001 — the record preserves the trajectory of how the claim was tested and failed, not merely the verdict that it failed — and remains reopenable only upon a future qualifying event (OQ-002).
  4. 2026-09-08 Collapsed LPR-001-D10 corrects several historical overstatements without changing the record's terminal judgement. The 1989 DOE review did not prove every anomalous heat report to be an artefact; rather, it found no convincing association between reported heat and a nuclear process, found the evidence for a new cold-fusion process unpersuasive, and highlighted severe reproducibility and fusion-product inconsistencies. The 2004 DOE review likewise did not produce a positive reversal: reviewer views on excess power were mixed, while most reviewers did not find the evidence for low-energy nuclear reactions conclusive and none recommended a focused federal programme. Berlinguette et al. (2019), correctly attributed here, then conducted a modern multi-institution re-evaluation that yielded no evidence of the cold-fusion effect. Continued LENR research and unresolved anomalous-effect claims remain historically relevant, but they do not supply reproducible evidence that an electrochemical cell itself produces nuclear fusion at or near room temperature. COLLAPSED / VS-05 is retained on this narrower, source-faithful basis.
  5. 2026-09-08 Collapsed Normal Record Review admits IN-008 and IN-009 as materially relevant new evidence at the boundary of the cold-fusion claim. The 2025 Nature result establishes that electrochemical deuterium loading can reproducibly increase the rate of D–D fusion in palladium when fusion is already being driven by externally accelerated deuterium ions. The 2026 Nature Communications result goes further mechanistically, showing a reproducible sub-keV fusion-yield plateau in electrochemically loaded palladium and titanium hydrides and a very large enhancement over bare-nucleus expectations, indicating that the condensed-matter environment can materially reshape low-energy fusion probabilities. These results are scientifically important and directly rehabilitate part of the broader question that survived the 2019 Google programme: materials can influence low-energy nuclear reaction rates in ways that merit study. They do not, however, reproduce the canonical claim recorded here. In both experiments an external ion beam supplies the kinetic energy that initiates fusion; electrochemistry loads or modifies the target rather than independently producing nuclear fusion at room-temperature chemical energies. The evidential boundary is therefore sharper, not weaker: condensed-matter-assisted beam-driven fusion is now positively demonstrated, while autonomous fusion generated by an electrochemical cell remains unconfirmed. COLLAPSED / VS-05 is retained. Reopening would require evidence that crosses that boundary rather than evidence of externally driven fusion enhanced by electrochemical loading.
Read complete Frontier Record
FR-AM-0002 - Anomalous Excess Heat — Electrochemical Cells Beyond Conventional Chemistry
  1. 2024-01-15 Fragmenting The claim occupies an unusual position in the corpus. The anomalous heat observations have been made by credentialled researchers using dedicated calorimetric equipment over thirty-five years. They have not been definitively refuted — no study has demonstrated that the reported observations are entirely attributable to measurement error, and the Berliner et al. (2019) study explicitly declined to make that claim. At the same time, the observations have not been reproduced on demand by independent laboratories following a shared, agreed protocol. The strongest quantitative claim in the record — SRI International's loading-fraction correlation (INST-002) — has not been independently confirmed under fully controlled conditions, and the Storms preparation protocols (INST-003) that claim to improve reproducibility have not been validated against a rigorous baseline. The pressure state is FRAGMENTING: the phenomenon has neither been confirmed as real nor definitively attributed to measurement artefact, and no agreed controlled protocol yet exists that would let a null result be accepted as meaningful (BN-001).
  2. 2026-09-09 Fragmenting LPR-001-D11 materially narrows the evidence for anomalous electrochemical excess heat without resolving the claim. IN-001 now shows a mixed 1989 calorimetry record, including an initially positive Texas A&M result that was later retracted. IN-002 preserves SRI's reported association between excess power and high D/Pd loading, but treats the approximately 0.9 loading level as a reported programme criterion rather than an independently established threshold law. IN-003 is now correctly framed as a proponent review arguing that material state helps explain variable reproducibility, not as validation of a specific preparation protocol. IN-004 removes the strongest legacy overstatement: Berlinguette et al. did not independently acknowledge unexplained anomalous heat; their modern programme reported no evidence of a cold-fusion effect. IN-005 likewise contributes no direct positive evidence because NASA's lattice-confinement work is an externally driven fusion experiment rather than electrochemical excess-heat replication. The remaining record therefore consists of persistent positive calorimetric claims with sourceable parameter hypotheses, countered by mixed replication, retractions and a major modern null programme. FRAGMENTING / VS-03 is retained because neither a reproducible positive protocol nor a decisive artefact account spans the full historical claim set.
Read complete Frontier Record
FR-AM-0003 - Cuprate Superconductivity — Mechanism Identification
  1. 2024-01-15 Fragmenting The mechanism responsible for cuprate superconductivity has not been identified in the sense the claim requires. After nearly four decades of intensive research, the field possesses several well-developed theoretical frameworks — spin fluctuation models, RVB and related strongly-correlated electron theories, charge density wave coupling proposals — none of which has achieved sufficient community consensus, predictive completeness, or experimental confirmation to constitute identification. The 2015 Keimer et al. review formally acknowledged that no single theory accounts for all cuprate phenomenology, and that conclusion has not been overturned by subsequent work. The pressure state is FRAGMENTING: this is not fragmentation from diverging evidence across domains, but from genuine theoretical plurality — multiple frameworks that are each partially correct and none of which has been falsified or achieved consensus (BN-001). Quantum simulation of the Hubbard model (AT-001) is the clearest visible resolution path, though it has not yet been executed at a scale sufficient to settle the question.
Read complete Frontier Record
FR-AM-0004 - Commercial Fusion Power — Net Electricity at Grid Scale
  1. 2024-01-15 Escalating The claim requires three thresholds to be met simultaneously: net electricity at plant level, grid-scale capacity, and commercial viability. None has been demonstrated. The furthest-reached threshold is threshold 1 (net electricity), which has been approached but not achieved at the plant level — NIF achieved Q > 1 at target level, not at facility level. Thresholds 2 and 3 are not yet addressable by current experimental evidence. The pressure state is ESCALATING. The NIF ignition result (INST-002) demonstrates that positive fusion energy gain is achievable in the laboratory, a necessary, though not sufficient, precondition for all three thresholds even though it satisfies none of them directly. Substantial private capital (INST-003) and a public ITER/DEMO roadmap (INST-004) indicate the engineering path is being actively pursued, but threshold 1 (plant-level net electricity) remains undemonstrated, and thresholds 2 and 3 cannot yet be meaningfully assessed given the sequential dependency between them (BN-001).
  2. 2026-06-29 Escalating No threshold has been crossed since AS-001. Threshold 1 (plant-level net electricity) remains undemonstrated; SPARC's own net-energy target is dated for 2026 and is not yet realised as of this assessment. What has changed is the density of engineering-milestone activity: SPARC assembly beginning, General Fusion's first-plasma result, and a coordinated DOE commercialisation roadmap all occurred within roughly the same window (late 2025), constituting the most concentrated burst of public engineering progress since the 2022 NIF/JET results that originally moved this record into ESCALATING. None of IN-006's events individually changes the assessment — they are pre-threshold engineering progress, the same evidence category as IN-003 — but their concentration is itself worth noting against OQ-001's resolution-criteria question: if SPARC's stated 2026 net-energy target is met, the Observatory will need exactly the governed procedure OQ-001 asks for and does not yet have.
Read complete Frontier Record
FR-AM-0005 - Room-Temperature Superconductivity — Reproducibility Under Laboratory Conditions
  1. 2024-01-15 Collapsed The claim has not been satisfied. No room-temperature superconductor has been reproduced under independent laboratory conditions to the community's current evidence standards. The two most prominent recent claims (Dias, LK-99) both failed replication — one through misconduct findings, one through rapid systematic null results from over forty independent groups. The confirmed high-pressure hydride results (INST-003) demonstrate that reproducible superconductivity approaching room temperature is achievable under extreme pressure, but not at room temperature or ambient pressure, and no material has met the community's evidence standard — zero resistance, Meissner effect, and specific heat anomaly, all independently confirmed — at conditions resembling laboratory practicality. The pressure state is COLLAPSED: the claim has been tested repeatedly, most recently and most rapidly in the LK-99 episode (INST-002), and no candidate has survived independent replication. The record remains open to reopening under AT-001 should a future material meet the tightened standard.
  2. 2026-06-29 Collapsed The claim remains unsatisfied and the pressure state remains COLLAPSED. IN-006 documents that the field did not go quiet after the 2024 null result — a nickelate stabilisation at ambient pressure (Feb 2025), a new 298K high-pressure record (Nov 2025, unreplicated), and a March 2026 field-wide research roadmap all represent real activity — but none meets AT-001's reopening condition: zero resistance, Meissner effect, and specific-heat anomaly, confirmed independently, under the community's tightened standard. The November 2025 result is the closest superficial match to a 'room-temperature' headline since LK-99, and is explicitly logged here so that the record does not appear to have missed it; on examination it fails the same threshold IN-001 through IN-003 already established — high pressure, no independent confirmation, no full evidentiary set. This assessment exists to confirm the COLLAPSED state remains correct under current evidence, not to revise it. The record's status remains CLOSED.
Read complete Frontier Record
FR-AM-0006 - Solid-State Batteries — Commercial Viability for Electric Vehicles
  1. 2024-01-15 Escalating The claim has not been satisfied. No solid-state battery has simultaneously demonstrated commercially viable energy density, safety, and cycle life at the manufacturing scale and cost required for EV deployment. Individual thresholds have been approached or met in laboratory settings; the three-threshold conjunction at commercial scale has not. The pressure state is ESCALATING. The field is advancing on genuine engineering problems with substantial industrial investment. The physics is not disputed — ion conduction through solid electrolytes is well understood — and the challenge is engineering and manufacturing scale-up rather than contested science (RM-001): laboratory milestones continue to be met and industrial commitment continues to grow (INST-002, INST-005), but the three-threshold conjunction at commercial manufacturing scale and cost has not been demonstrated, and announced delivery timelines have consistently receded rather than been met (INST-003).
  2. 2026-06-27 Escalating INST-006 sustains the ESCALATING state identified at AS-001 rather than advancing or collapsing it. The 2025 commercial target already flagged as superseded at IN-003 has now genuinely elapsed without a solid-state EV reaching production, which removes any ambiguity about whether that particular date might still be met. At the same time, the evidence does not support reclassifying this record toward the PROG-AM collapse dynamic (CM-001 elsewhere in the corpus): government production approval for the underlying technology, a named material-supply joint venture with a defined 2027 facility start date, and continuing — if uneven — progress from Chinese manufacturers are all genuine engineering and industrial advances, not disputed physics or failed replication. The pattern remains exactly what RM-001 describes: laboratory and component-level milestones continue to be met while full commercial-scale, three-threshold delivery continues to recede. Verification stage advances to VS-03 (Audit): the underlying technology has now cleared a formal government regulatory/production-approval review, the first independent scrutiny event in this record's history — though this is approval of the technology rather than independent replication of Toyota's specific performance claims (IN-004), which remains unverified in peer-reviewed form. OQ-002's procedural question (whether dated attractors warrant scheduled re-entry) is now reinforced by direct example: this record's own dated attractor target has elapsed.
  3. 2026-08-29 Escalating IN-007 sustains ESCALATING / VS-03. QuantumScape's automated Eagle Line, customer sample shipments, and milestone-based PowerCo programme are stronger evidence of industrialization than the earlier single-layer and prototype milestones in this record. They directly bear on RM-001 because the work is now testing repeatable manufacturing processes rather than only electrochemical performance. But the same primary filings explicitly preserve the unresolved commercial gap: QuantumScape remains pre-revenue, the line is still a pilot facility, and quality, consistency, reliability, throughput, safety, and cost remain development requirements. No evidence reviewed in this pass demonstrates a production EV battery meeting energy density, safety, and cycle-life requirements simultaneously at commercial manufacturing yield and cost. The attractor therefore remains future operational deployment rather than pilot-line progress.
  4. 2026-08-29 Escalating Classification correction following bounded source review. ESCALATING is retained, but the current Verification Stage returns from VS-03 to VS-02. AS-002 advanced the record to VS-03 because a Japanese government event was characterised as a regulatory/production-approval review providing independent scrutiny of the technology. The underlying event was instead METI certification, on September 6, 2024, of Toyota’s battery development and production plan under the Battery Supply Assurance Plan. That industrial-policy certification supports the reality and seriousness of Toyota’s programme, but it does not independently audit Toyota’s claimed solid-state battery performance, manufacturing yield, safety, cycle life, energy density, or commercial viability. The later Sumitomo Metal Mining agreement and the 2026 QuantumScape pilot-line evidence likewise strengthen the industrialisation trajectory without supplying the independent claim-level scrutiny required for VS-03. This stage correction is epistemic: it corrects the Observatory’s earlier classification rationale and does not represent deterioration in the technology or a change in the ESCALATING pressure state.
Read complete Frontier Record
FR-AM-0007 - Pressure-Quenched Superconductivity — Retention of High-Pressure States at Ambient Pressure
  1. 2026-08-25 Escalating The claim is supported by a cumulative experimental trajectory rather than a single headline result. Pressure-quench retention has been reported across multiple superconducting materials, culminating in the 2026 Hg1223 result retaining an enhanced transition temperature up to 151 K after decompression. That progression is sufficient to move the claim beyond EMERGING: the phenomenon has recurred across material systems and has been subjected to peer-reviewed experimental characterisation. The pressure state is ESCALATING because the evidence base is expanding in strength and generality while the decisive uncertainties remain open. The principal unresolved issue is independent replication outside the originating research network (RM-002). A second limitation is physical durability: ambient pressure is not equivalent to ambient-condition stability, because the retained Hg1223 state is metastable and degrades on warming (RM-001). Verification Stage is VS-03 — Audit: the published evidence has substantial internal controls and cross-material recurrence, but no unaffiliated laboratory has yet reproduced the pressure-quench effect under a shared protocol. Independent replication (AT-001) is therefore the next evidential boundary.
Read complete Frontier Record
FR-BT-0001 - Senolytic Therapies — Meaningful Human Healthspan Extension
  1. 2024-01-15 Escalating The claim has not been satisfied. No senolytic therapy has demonstrated meaningful healthspan extension in humans on clinical endpoints. The foundational preclinical evidence (INST-001) establishes a compelling causal mechanism — senescent cell accumulation contributes to aging, and their removal produces healthspan benefit in mice. The human surrogate evidence (INST-002, INST-004) demonstrates that senolytics reduce senescent cell burden in humans. But the Phase II clinical trial failures (INST-003) — the first adequately powered randomised trials of senolytics in humans — failed to demonstrate benefit on primary clinical endpoints. The pressure state is ESCALATING: the mechanistic and surrogate-marker case remains strong, and the first hard clinical test has returned a null result that is attributable at least partly to drug choice, dosing, and endpoint selection (RM-002) rather than a clean refutation of the underlying hypothesis, but the surrogate-to-clinical translation gap (RM-001) is now the central unresolved obstacle.
Read complete Frontier Record
FR-BT-0002 - Epigenetic Reprogramming — Biological Age Reversal Without Identity Loss
  1. 2024-01-15 Escalating The claim has not been satisfied in humans. Partial epigenetic reprogramming without loss of cellular identity has been demonstrated in multiple mouse models and is extending toward non-human primates. No human clinical trials have been initiated. The biological mechanism is well-established: OSKM and related factors can reset epigenetic age marks; partial expression can do so without completing dedifferentiation; and the process produces functional improvements in at least some mouse tissues. The claim's human clinical evidence gap remains complete: no partial reprogramming therapy has yet entered a human trial. The pressure state is ESCALATING: the mechanism is well established across multiple mouse models and the field is heavily capitalised (INST-003), but whether partial reprogramming is safe and effective in humans — and whether epigenetic clock reversal constitutes genuine rejuvenation rather than a movable measurement (BN-001) — remains entirely untested outside model organisms.
  2. 2026-06-29 Escalating The human clinical evidence gap that AS-001 identified as complete is now closing. Life Biosciences has received FDA IND clearance for ER-100, a partial OSK reprogramming therapy, with a stated trial start of Q1 2026 — the first human trial of any partial epigenetic reprogramming therapy. This is the first half of AT-001's named resolution attractor ("first human safety data and validated functional outcome biomarkers"); the second half — actual safety and clock-reversal data — does not yet exist, since the trial has only just been cleared to begin, not completed or reported. The pressure state remains ESCALATING rather than moving to RESOLVING: clearance to run a trial is a regulatory and operational milestone, not efficacy or safety evidence. BN-001 (clock validity as a rejuvenation surrogate) is unaffected by this development and remains the record's primary interior bottleneck regardless of how the ER-100 trial proceeds. This assessment exists to record that the record's own named attractor condition has begun to materialise, not to anticipate its outcome.
  3. 2026-08-29 Escalating PA-006 provenance-in-review final replication confirms that ER-100 has progressed from regulatory clearance to actual human dosing: Life Biosciences reported the first participant dosed on June 9, 2026, and ClinicalTrials.gov lists NCT07290244 as recruiting. This is a substantive operational advance because the claim is now being tested directly in humans rather than only authorised for testing. It does not yet satisfy AT-001 or the governing claim. No results are posted, so there is still no human evidence establishing safety at partial-reprogramming doses, biological-age reversal, preserved cellular identity, or validated functional rejuvenation. BN-001 therefore remains unresolved. ESCALATING / VS-02 is retained pending human outcome evidence.
Read complete Frontier Record
FR-BT-0003 - Biological Age Biomarker Panels — Predictive Validity for Age-Related Decline
  1. 2024-01-15 Fragmenting The claim is partially supported and fragmenting. Blood-based biomarker panels demonstrate population-level predictive validity for biological aging outcomes — at the population level, high biological age scores predict faster subsequent decline, higher mortality risk, and earlier disease onset. This is well-established across multiple panel types (epigenetic, proteomic, metabolomic) and multiple longitudinal cohorts. The population-level claim is supported. The claim fragments at the individual level. Different biological age clocks give substantially different estimates for the same individual, and organ systems within one person age at markedly different rates (INST-003): a blood panel captures a composite population-level signal that may not reflect which specific organ or process is declining fastest in any given person. The pressure state is FRAGMENTING: population-level predictive validity is well established, but individual-level predictive validity — the form the claim requires for clinical use — has not been demonstrated, and the actionability gap identified in DunedinPACE (INST-004) means that even a valid individual-level signal may not yet translate into a clear intervention (BN-001).
Read complete Frontier Record
FR-BT-0004 - Liquid Biopsy — Early Cancer Detection Before Conventional Diagnosis
  1. 2024-01-15 Fragmenting The claim is partially supported and fragmenting. Blood-based liquid biopsy can detect cancer signals before conventional diagnosis in a demonstrable fraction of cases — the Galleri test's performance data establishes this for multiple cancer types. The detection is reliable in a technical sense: specificity is high (98.4%) and sensitivity, while lower than desired, is non-trivial across cancer types. For the claim as stated, this constitutes partial confirmation: early detection before conventional diagnosis is achievable in some cases. The claim fragments on the central unresolved question: whether that earlier detection reduces cancer mortality, or whether it produces stage shift and lead-time effects without a genuine survival benefit (IN-004). Sensitivity is markedly lower for early-stage disease (approximately 24% at Stage I) — precisely the regime in which the claim's value would be greatest — and the NHS-Galleri trial (INST-003), the first to test the claim against a mortality endpoint directly, has not yet reported. The pressure state is FRAGMENTING: the technology works as a detection instrument, but whether detection translates into the clinical benefit the claim asserts remains genuinely open (AT-001).
  2. 2026-06-27 Fragmenting The NHS-Galleri trial's full results (INST-006) sustain rather than resolve the FRAGMENTING state identified at AS-001. The trial delivers exactly the kind of evidence the record's attractor (AT-001) was built to await, and the result is genuinely mixed rather than confirmatory or disconfirming: a real, substantial reduction in late-stage diagnoses coexists with a missed primary endpoint, an unexpected rise in Stage III diagnoses, and no mortality data. This is not a null result — the four-fold detection-rate increase and Stage IV reduction are real signals — but it does not resolve the central contested question (OQ-001): whether earlier detection translates into reduced mortality, or whether it is partially absorbed by stage migration and lead-time effects that RM-001/AT-001 already anticipated. Verification stage advances to VS-04 (Replication): a population-scale randomised trial has now run and reported, the most rigorous test design available short of mortality follow-up itself. The record should be re-entered when GRAIL's extended follow-up data (6–12 months from this release) becomes available, since that data — not this release — is positioned to address OQ-001 directly.
Read complete Frontier Record
FR-BT-0005 - Gene-Edited Porcine Kidneys — Durable Human Renal Replacement
  1. 2026-08-21 Escalating The claim enters the corpus under strong two-sided pressure. Living-human xenokidney recipients have demonstrated life-sustaining renal function for weeks and, in a later case, for 271 days, moving the field beyond decedent compatibility and short-lived physiological demonstration. At the same time, acute rejection, persistent innate immune activation, eventual graft dysfunction, proteinuria and the continuing need for intensive immunosuppression and zoonotic surveillance remain material constraints. The EXPAND prospective multicentre study now provides a formal replication pathway with 24-week graft, patient and renal-function endpoints. The Pressure State is ESCALATING because both capability evidence and failure-mechanism evidence are strengthening. Verification Stage is VS-03 because the claim has progressed beyond publication into living-human clinical audit that has exposed real rejection and physiological failure modes, but prospective multi-recipient replication has not yet been established.
Read complete Frontier Record