Faultline Observatory

Evidence Trajectories

Follow how the Observatory's judgement of technology claims has changed as evidence accumulated.

Each line represents one Frontier Record. Select a reading to bring related trajectories forward; the full archive remains present.

How to read this page
  1. Read across - time moves from left to right.
  2. Read vertically - position indicates verification depth at that moment.
  3. Read the register - the right column names each record's current Verification Stage at Today.
  4. Open the record - every trajectory can be traced back to its documentary history.

Verification-stage provenance notice: the legacy review is complete. Historical codes remain preserved; trajectories apply the ratified evidence-depth review overlay. Assignments that could not be reliably reconstructed remain marked historically unverified.

Evidence TrajectoriesEvidence Trajectories arranges documented assessments across four evidence-history phases. Horizontal spacing represents documentary sequence rather than equal calendar time. The Current Register is grouped by current verification stage.OperationVS-05AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined TasksFR-AI-0008AI Medical Imaging DiagnosisReplicationVS-04Cold Fusion — Room-Temperature Nuclear Fusion in Electrochemical CellsFR-AM-0001Cold FusionAnomalous Excess Heat — Electrochemical Cells Beyond Conventional ChemistryFR-AM-0002Anomalous Excess HeatRoom-Temperature Superconductivity — Reproducibility Under Laboratory ConditionsFR-AM-0005Room-Temperature SuperconductivityBiological Age Biomarker Panels — Predictive Validity for Age-Related DeclineFR-BT-0003Biological Age Biomarker PanelsLiquid Biopsy — Early Cancer Detection Before Conventional DiagnosisFR-BT-0004Liquid BiopsyGoogle Quantum Advantage (Sycamore) — Computational Supremacy via Random Circuit SamplingFR-QE-0001Google Quantum Advantage (Sycamore)Below-Threshold Quantum Error Correction — Scalable ArchitectureFR-QE-0004Below-Threshold Quantum Error CorrectionQuantum Error Correction Scaling — Logical Rate Suppression Under Physical OverheadFR-QE-0008Quantum Error Correction ScalingAuditVS-03LLM Multi-Step Reasoning — Generalisation Beyond TrainingFR-AI-0001LLM Multi-Step ReasoningRLHF Preference Generalisation — Behaviour Beyond Training DistributionFR-AI-0003RLHF Preference GeneralisationScaling Laws — Emergent Performance on Unseen TasksFR-AI-0004Scaling LawsAGI Through Scaling — LLM Architecture as the Path to General IntelligenceFR-AI-0005AGI Through ScalingScaling Mechanism Coherence — Continuity Across Model SizesFR-AI-0006Scaling Mechanism CoherenceAutonomous AI Scientific Discovery — Novel, Correct, IndependentFR-AI-0007Autonomous AI Scientific DiscoveryCuprate Superconductivity — Mechanism IdentificationFR-AM-0003Cuprate SuperconductivityCommercial Fusion Power — Net Electricity at Grid ScaleFR-AM-0004Commercial Fusion PowerSolid-State Batteries — Commercial Viability for Electric VehiclesFR-AM-0006Solid-State BatteriesSenolytic Therapies — Meaningful Human Healthspan ExtensionFR-BT-0001Senolytic TherapiesEpigenetic Reprogramming — Biological Age Reversal Without Identity LossFR-BT-0002Epigenetic ReprogrammingD-Wave Quantum Annealing — Practical Computational AdvantageFR-QE-0002D-Wave Quantum AnnealingFault-Tolerant Logical Qubits — Error Rate Scaling with Code DistanceFR-QE-0003Fault-Tolerant Logical QubitsCryptographically Relevant Quantum Computing — RSA FactorisationFR-QE-0005Cryptographically Relevant Quantum ComputingPractical Quantum Advantage — Performance Beyond Classical Computation on Relevant ProblemsFR-QE-0007Practical Quantum AdvantagePublishedVS-02LLM Knowledge-Work Utility — Economically Valuable Task PerformanceFR-AI-0002LLM Knowledge-Work UtilityFault-Tolerant Quantum Utility — Practically Useful Algorithms Beyond Classical SimulationFR-QE-0006Fault-Tolerant Quantum UtilityAssertionVS-01No current records

The trajectories above are derived from the assessment histories below.

Trajectory records

SHOWING THE FULL ARCHIVE · NO READING IS PRESELECTED.

FR-QE-0001Google Quantum Advantage (Sycamore) — Computational Supremacy via Random Circuit SamplingA programmable quantum processor has demonstrated computational supremacy — performing a well-defined sampling task beyond the practical reach of any classical computer.QE · 2 assessments · 1 documented state change
Stabilisingsince 2026-07-22
Assessment history
  1. 2026-06-11Fragmenting

    The original 2019 quantum supremacy claim has undergone a complete evidence cycle. At announcement, the claim was precise and measurable: 200 seconds versus an estimated 10,000 classical years on a specific Random Circuit Sampling task. Within five years, classical simulation methods improved by orders of magnitude, culminating in the Zhao et al. (2024) result demonstrating classical performance exceeding the original quantum benchmark in both speed and fidelity. The original claim, as stated in 2019, has been effectively superseded. However, the research programme that produced the claim has not collapsed. Google's Willow processor (December 2024) reasserts the supremacy framing at a vastly larger scale (10²⁵ classical years) while simultaneously demonstrating below-threshold quantum error correction — a qualitatively different and more durable achievement. The evidence trajectory has therefore split: the narrow 2019 benchmark claim is contested to the point of supersession, while the broader programme claim (that quantum processors are advancing toward practical computational advantage) has arguably strengthened. This record exhibits the canonical pattern the FCIF subsequently formalised as Claim Migration: the original claim does not resolve cleanly (neither fully vindicated nor retracted) but instead evolves as the claimant shifts the evidential basis to a new formulation that inherits the original's ambition but rests on different technical foundations. The 2019 claim migrated from "supremacy via RCS on Sycamore" to "scalable error correction via surface codes on Willow." The Observatory records this as a Fragmenting state: the evidence does not converge on a single verdict because the claim itself has moved.

  2. 2026-07-22Stabilising

    Institutional verdict: UNRESOLVED AT THE ORIGINAL COMPARISON POINT; LATER SUPERSEDED. The admitted claim is a time-indexed comparative-performance claim concerning Sycamore's 2019 Random Circuit Sampling demonstration. The experiment and task performance are supported, but the constitutive claim that the task was beyond practical classical reach was not established at the original comparison point because IBM's contemporaneous days-scale analysis left that threshold unresolved. Later classical work, culminating in Zhao et al. (2024), reproduced and surpassed the benchmark comparator. That later result supersedes the demonstrated advantage without automatically proving that the time-indexed 2019 claim was false when made. Pressure State is STABILISING. The uncertainty is now bounded and durable rather than fragmenting: Willow and other successor claims are outside this claim's material commitments, and the original comparison can reopen only through evidence showing that the 2019 comparator was already unsound. Verification Stage is VS-04 — Replication because independent classical work progressed beyond audit to direct reproduction and eventual counter-performance of the claim's constitutive comparator. This is claim-level, adversarial replication; it does not assert independent reproduction of Sycamore hardware or favourable confirmation of the original advantage.

FR-QE-0002D-Wave Quantum Annealing — Practical Computational AdvantageQuantum annealing systems have demonstrated practical computational advantage over classical methods on commercially or scientifically relevant optimisation tasks.QE · 2 assessments · 0 documented state changes
Fragmentingsince 2026-07-26
Assessment history
  1. 2024-01-15Fragmenting

    The evidence trail for this claim is fragmented across distinct problem domains and claim interpretations. On commercially motivated optimisation tasks (scheduling, routing, combinatorial problems of practical scale), no published evidence has established durable advantage over state-of-the-art classical methods. The contested Denchev et al. (2016) result represents the strongest performance claim in this domain; it was substantially undermined by subsequent classical algorithm improvements and the benchmark's structural dependence on hardware-favourable problem instances (RM-002). In the separate domain of scientific simulation, the evidence is stronger: King et al. (2022, 2023) report computational advantage in simulating quantum magnetism, though critics dispute the comparison class used. The claim spans two domains accruing evidence asymmetrically and has not been decomposed into separate records (BN-001), and no agreed classical comparison class exists (BN-002). The pressure state is FRAGMENTING: the claim is not converging toward a single assessment but splitting along domain lines that may require separate evaluation.

  2. 2026-07-26Fragmenting

    The Pressure State is unchanged. Its governing rationale is not. The prior domain-split explanation (AS-001) is retired following the ratified Identity and Continuity Review, which found IDENTITY PRESERVED AS A COMPOUND CLAIM: a single recoverable kernel — quantum annealing, practical computational advantage, a classical comparator, optimisation-task class, commercial-or-scientific relevance — is engaged by evidence from both relevance routes. IN-004 is adjacent simulation evidence and does not bear on this claim. Among the remaining instances, IN-001 through IN-003 read negative-to-contested on the commercial route across eight years, and IN-005 provides a single, contemporaneously-grounded positive instance whose own comparator (quantum Monte Carlo) is disputed (BN-002). FRAGMENTING is warranted not because the claim splits along commercial or scientific lines, but because this same kernel-corrected evidence supports incompatible trajectory interpretations under unresolved competing meanings of 'practical advantage' — whether that standard requires real-world deployability or a rigorous demonstration of speedup on a well-posed instance (OQ-6, proposed). This is interpretive, not referential, non-convergence: no identity fracture, no decomposition, no admission-scope defect.

FR-QE-0003Fault-Tolerant Logical Qubits — Error Rate Scaling with Code DistanceFault-tolerant logical qubits can be demonstrated with logical error rates that improve as error-correcting code distance increases.QE · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-28
Assessment history
  1. 2024-01-15Escalating

    The claim describes a specific empirical signature: logical error rates improving as code distance increases. This signature has now been demonstrated. INST-003 (Google, 2023) was the first result to show simultaneous X and Z error suppression with increasing code distance, directly satisfying the claim's measurement criterion. INST-005 (Google Willow, 2024) extends this to below-threshold operation, showing that the improvement rate exceeds the overhead rate — the condition required for the result to be considered scalable rather than merely demonstrated at fixed size. INST-001 (2021) and INST-002 (2022) provided earlier partial evidence — single-error-type suppression and below-physical-error-rate operation, respectively — establishing the trajectory that INST-003 and INST-005 confirm more directly. The pressure state is ESCALATING: the claim's core empirical signature is demonstrated and strengthening, but the demonstrated code distances (up to 7) remain well below the distances required to confirm the behaviour holds at practically relevant scale (OQ-001).

  2. 2026-06-28Escalating

    The evidence gap is closed by IN-006. The Willow result is now treated as a verified, peer-reviewed below-threshold surface-code memory result rather than a general quantum-computing announcement. It materially strengthens the claim because logical error suppression improves with code distance and the larger memory exceeds break-even. The pressure state remains ESCALATING rather than RESOLVING because the record's own next decisive question — whether below-threshold scaling holds at d=11 and above — remains unanswered, and the demonstrated result is still a memory result rather than a full fault-tolerant computation pathway.

FR-QE-0004Below-Threshold Quantum Error Correction — Scalable ArchitectureQuantum error correction can reduce logical error rates below physical error rates in a scalable architecture.QE · 1 assessments · 0 documented state changes
Resolvingsince 2024-01-15
Assessment history
  1. 2024-01-15Resolving

    The claim has two components: below-physical-rate operation, and scalability of that operation. Both have been demonstrated. INST-002 established that below-physical-rate logical qubits are achievable in principle. INST-003 established that practically useful error rates are achievable on current hardware. INST-004 established that performance improves as the architecture scales — the defining signature of scalable below-threshold operation. No contesting evidence has been published against either component of the claim. The pressure state is RESOLVING: both elements the claim requires have been independently demonstrated and corroborate each other, though confirmation at the code distances required for practical fault-tolerant computation (d=11 and above) remains outstanding (OQ-001), and the claim's architecture-agnosticism has not yet been confirmed across a third hardware platform (OQ-002).

FR-QE-0005Cryptographically Relevant Quantum Computing — RSA FactorisationA quantum computer can factor commercially relevant RSA cryptographic keys faster than any classical computer.QE · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-29
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No quantum computer has factored a commercially relevant RSA key. The most credible direct attempt (INST-005) failed. The engineering gap between current capability and the Gidney-Ekerå resource estimate remains approximately three to four orders of magnitude in physical qubit count, with additional requirements for error rates, connectivity, and operational duration not yet demonstrated at any scale approaching relevance. The pressure state is ESCALATING rather than EMERGING because the substrate advances documented in FR-QE-0003 and FR-QE-0004 (INST-003) show the underlying error-correction engineering progressing on a credible trajectory, even though the gap to the resource requirement remains enormous. Institutional behaviour — NIST's finalisation of post-quantum cryptography standards (INST-004) — reflects institutional acceptance that the risk is credible enough to justify migration, adding pressure to the claim's trajectory independent of any direct technical progress toward satisfaction.

  2. 2026-06-29Escalating

    No threshold has been crossed since AS-001 — no factorisation of a commercially relevant key has occurred, and none is closer to occurring in any demonstrated sense. What has moved is the resource-estimate trajectory underlying OQ-001. Gidney (Google, May 2025) reduced the estimated physical-qubit requirement for RSA-2048 factorisation from the Gidney-Ekerå (2021) figure of ~20 million to under 1 million, under comparable fault-tolerance assumptions — roughly a 20-fold reduction achieved through improved algorithmic and error-correction engineering rather than any experimental demonstration. A 2026 proposal using QLDPC codes (an architecture distinct from the surface codes assumed in both prior estimates) suggests a further reduction toward ~100,000 physical qubits, though this is unvalidated at scale. A March 2026 Google/Stanford/Ethereum Foundation whitepaper applies the same style of resource-reduction analysis to elliptic-curve cryptography, estimating under 500,000 physical qubits for widely used curves. All three results are theoretical resource estimates — the same evidence category as INST-002's original figure — not experimental progress toward the claim. The pressure state remains ESCALATING; no reclassification is warranted by an estimate revision alone. What is new is the rate: three independent downward revisions within roughly eighteen months is faster compression of the engineering-gap estimate than the original record anticipated, and OQ-001 now has materially fresher input than it did at AS-001.

FR-QE-0006Fault-Tolerant Quantum Utility — Practically Useful Algorithms Beyond Classical SimulationA fault-tolerant quantum computer can execute a practically useful quantum algorithm beyond classical simulation.QE · 1 assessments · 0 documented state changes
Escalatingsince 2024-01-15
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No fault-tolerant quantum computer has executed a practically useful quantum algorithm beyond classical simulation at the scale required for genuine practical advantage. The substrate progress (INST-003) establishes that fault-tolerant logical qubits capable of executing simple circuits now exist; the resource estimation (INST-004) establishes that practically useful chemistry simulation requires approximately two orders of magnitude more logical qubits than are currently available. The gap is smaller and more tractable than the equivalent gap for RSA factorisation (FR-QE-0005), but still represents years of further engineering. Classical simulation methods are simultaneously improving (INST-005), narrowing the space of problems that would unambiguously qualify as beyond classical reach by the time fault-tolerant hardware reaches the required scale. The pressure state is ESCALATING: the substrate is advancing on a credible path, but no agreed target problem yet exists (BN-001) on which the claim could be tested.

FR-QE-0007Practical Quantum Advantage — Performance Beyond Classical Computation on Relevant ProblemsA quantum computer has achieved quantum advantage on a practically relevant problem.QE · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The claim has not been satisfied. No quantum computer has demonstrated advantage on a problem that simultaneously meets both the performance threshold (faster than best classical methods) and the practical relevance threshold (problem has genuine scientific or commercial value at the demonstrated scale). The evidence base contains strong demonstrations of one component without the other — advantage on demonstration problems (INST-001, 002, 004) or near-advantage on relevant problems (INST-005) — but no instance yet satisfies both simultaneously. IBM's quantum utility claim (INST-003) comes closest to bridging the two, reporting results on a problem with some scientific relevance that classical simulation was disputed to match, but the classical-simulation contest remains unresolved. The pressure state is FRAGMENTING: the evidence is splitting along two separate trajectories — demonstration-problem advantage growing stronger (INST-004) and relevant-problem simulation approaching but not reaching classical intractability (INST-005) — without converging on a single instance that would resolve the claim (OQ-001).

FR-QE-0008Quantum Error Correction Scaling — Logical Rate Suppression Under Physical OverheadQuantum error correction can reduce logical error rates faster than physical error rates increase with system scale.QE · 1 assessments · 0 documented state changes
Resolvingsince 2024-01-15
Assessment history
  1. 2024-01-15Resolving

    The claim is substantially supported and on a trajectory toward confirmation. Google's Willow results (INST-003) demonstrate exponential logical error rate suppression through code distance 7, consistent with the threshold theorem's predictions. Cross-platform confirmation from Microsoft and Quantinuum (INST-004) strengthens the result beyond a single-platform observation. The core scaling relationship — logical error rates suppressing faster than physical overhead increases — is empirically confirmed at the code distances tested. The pressure state is RESOLVING: the theorem's central prediction has been consistently observed across the code distances measured so far (INST-002, INST-003) and across multiple hardware architectures, though correlated-error effects that may limit suppression at larger code distances (INST-005) have not yet been ruled out, and confirmation at the distances required for practical fault tolerance (d=9, d=11) remains the decisive open step (AT-001).

FR-AI-0001LLM Multi-Step Reasoning — Generalisation Beyond TrainingLarge language models can perform multi-step reasoning that generalises beyond memorised training examples.AI · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-27
Assessment history
  1. 2024-01-15Escalating

    The evidence trail shows a claim under genuine escalating pressure. Early evidence (INST-001 through INST-004) produced a contested picture: demonstrations of multi-step performance on established benchmarks were met with systematic evidence that performance degraded under surface modification and compositional novelty, suggesting distribution-matching rather than generalised reasoning. That picture was the dominant assessment context through 2023. INST-005 materially shifts the evidentiary state. o3's 87.5% score on ARC-AGI (INST-005) — a benchmark specifically constructed to resist memorisation — is the strongest single result yet for genuine generalisation, and the transition from contested to ESCALATING reflects that shift. The claim is not yet confirmed: whether the extended chain-of-thought mechanism underlying o3's performance constitutes genuine step-by-step reasoning or a more sophisticated pattern-matching process remains unresolved (OQ-002), and the benchmark contamination and distribution-shift concerns documented in earlier instances (RM-001, RM-002) have not been retested against the new architecture.

  2. 2026-06-27Escalating

    INST-006 sustains the ESCALATING state from AS-001 while materially sharpening OQ-002 rather than closing it. The disclosure that chain-of-thought traces are frequently unfaithful — not reliably reflecting the computation that produced an answer — means o1/o3-class performance (INST-005) cannot be straightforwardly read as evidence of the reasoning process its own output narrates. This cuts against treating AT-001's mechanism candidate as settled in either direction: a model could be performing genuine multi-step computation that its verbalised trace merely fails to describe accurately, or could be pattern-matching while its trace fabricates a plausible reasoning narrative — the faithfulness literature establishes that both are observed, without yet establishing which dominates for any specific frontier system. The claim's evidentiary picture therefore escalates in complexity: BN-001's undefined generalisation threshold is now joined by an analogous undefined-faithfulness threshold, and OQ-002 should be read going forward as two distinct questions (does the model generalise; does its chain-of-thought narrate that generalisation faithfully) rather than one. Verification stage advances to VS-03 (Audit): a substantial, multi-author, partly first-party (Anthropic) literature has now subjected the mechanism itself to direct scrutiny — the first such audit-stage evidence this record has logged.

FR-AI-0002LLM Knowledge-Work Utility — Economically Valuable Task PerformanceLarge language models can perform economically valuable knowledge-work tasks with limited human supervision.AI · 1 assessments · 0 documented state changes
Escalatingsince 2024-01-15
Assessment history
  1. 2024-01-15Escalating

    The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).

FR-AI-0003RLHF Preference Generalisation — Behaviour Beyond Training DistributionReinforcement learning from human feedback produces AI systems whose behaviour continues to reflect human preferences when deployed beyond the conditions represented in training.AI · 2 assessments · 0 documented state changes
Fragmentingsince 2026-07-14
Assessment history
  1. 2024-01-15Fragmenting

    The evidence trail for this claim does not converge. Three distinct failure modes have been documented under three distinct kinds of distribution shift: adversarial prompting (INST-002), novel social context producing approval-seeking (INST-003), and capability gains that outpace preference calibration (INST-005). These are not the same mechanism and they are not reducible to each other. A system that solved the adversarial prompting problem would not automatically solve sycophancy; a system that solved sycophancy would not automatically be robust to capability-outpacing drift. Constitutional AI (INST-004) shows that training methodology improvements can partially address these failure modes without resolving them, and weak-to-strong generalisation research (INST-006) suggests these failure modes may not be structurally unavoidable even as capability increases outpace preference calibration (INST-005). The pressure state is FRAGMENTING: the claim's failure modes are documented but distinct, and no single mechanism or measurement approach yet unifies them (BN-001).

  2. 2026-07-14Fragmenting

    Pressure state FRAGMENTING is retained. What changed: the capability/generalisation tension identified in AS-001 as this record's central unresolved question (OQ-001) has received its first direct empirical pressure. ROGUE (IN-007) measures corrigibility failure under ordinary — not adversarial — deployment conditions and finds that better-performing models exhibit greater misalignment, the first empirical datapoint bearing directly on whether increasing capability makes generalisation worse; within the tested regime it points toward worse. Two independent sources corroborate a route by which deployed-agent behaviour may fail that is not cleanly captured by the existing three failure modes (adversarial IN-002, sycophantic IN-003, capability-outpacing IN-005): a structural argument that in-weights safety training does not transfer to agentic authority contexts (IN-008), and a bounded empirical finding of strategic public/off-record divergence under pressure (IN-009). What remains unresolved, and is the boundary this assessment records without deciding: whether this constitutes a fourth failure mode within the RLHF preference-generalisation claim, or a distinct agentic-corrigibility claim that warrants its own Frontier Record. The evidence deepens fragmentation; it does not resolve the claim in either direction. No corrigibility record is opened at this time — the class-level boundary question (cf. OQ-004 and the FR-QE-0002 over-bundling lesson) is left for further evidence to settle rather than pre-empted. The three new instances are contesting or bounded-contesting; none is a supportive convergence, and the FRAGMENTING state is sustained on that basis.

FR-AI-0004Scaling Laws — Emergent Performance on Unseen TasksScaling language model training increases performance on previously unseen tasks without task-specific optimisation.AI · 2 assessments · 0 documented state changes
Fragmentingsince 2026-06-29
Assessment history
  1. 2024-01-15Fragmenting

    The claim is supported in its core assertion: scaling language model training does increase performance on previously unseen tasks without task-specific optimisation. This is documented across multiple model families, task types, and evaluation methodologies. The few-shot performance documented in INST-002, the smooth scaling curves in INST-001 and INST-006, and the emergent task capabilities in INST-003 all constitute positive evidence for the claim as stated. The evidence trail is nonetheless complicated by two interior disputes that do not threaten the claim's truth but substantially complicate its mechanism: whether apparent emergent abilities (INST-003) are genuine discontinuities or artefacts of metric choice (INST-004), and whether benchmark performance gains reflect genuine generalisation or training-data contamination (INST-005). The pressure state is FRAGMENTING: the claim's core assertion holds, but the evidence quality disputes over how and why it holds have not converged, and no agreed definition of 'previously unseen' yet exists to resolve them (BN-001).

  2. 2026-06-29Fragmenting

    The claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.

FR-AI-0005AGI Through Scaling — LLM Architecture as the Path to General IntelligenceArtificial General Intelligence will be achieved through scaling current large-language-model architectures.AI · 2 assessments · 0 documented state changes
Fragmentingsince 2026-06-29
Assessment history
  1. 2024-01-15Fragmenting

    The claim is fragmenting in a structurally unusual way. The evidence trail shows neither clean positive progression nor clean negative accumulation. Instead it shows a claim under three simultaneous pressures that are each individually partial: capability gains continue (supportive), but the path is bifurcating architecturally (INST-004); structural scaling constraints are accumulating (INST-005); and the target itself is migrating (INST-006). These pressures do not converge on a single conclusion. Capability continues to advance in ways that keep the claim alive, while the path departs from pure scaling and the destination itself is redefined in ways that make the claim progressively harder to evaluate as originally stated. The pressure state is FRAGMENTING: the claim is not resolving toward confirmation or collapse but splitting along three independent axes — capability, path, and target — each of which would need to be separately addressed before the claim could reach a stable assessment (OQ-001).

  2. 2026-06-29Fragmenting

    FRAGMENTING remains the correct pressure state, and IN-007 is best read as confirmation rather than a new direction. The three simultaneous pressures AS-001 identified — capability gains continuing, the path bifurcating architecturally, the target migrating — have each continued through 2025–26 without converging. Field-wide commentary now describes 2026 progress as inference- and tooling-driven rather than training-scale-driven, and academic work documents diminishing (though not zero) returns on pure training-compute scaling. Simultaneously, frontier labs' continued tens-of-billions-dollar commitments to training-scale infrastructure through 2025 show the industry has not abandoned the original path either. No single development in IN-007 resolves OQ-001 (can a claim with a migrating target reach a stable assessment state) or OQ-002 (is this dissolution or collapse) — if anything, two more years of continued three-way fragmentation without resolution is itself mild evidence that this claim may be heading toward dissolution rather than either confirmation or collapse, which is exactly the distinction OQ-002 asks the Observatory to make a governance decision about.

FR-AI-0006Scaling Mechanism Coherence — Continuity Across Model SizesCapabilities that emerge through scaling language models are explained by the same underlying mechanism across model sizes.AI · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The evidence trail is genuinely mixed and the mixing is interior — it concerns what the mechanisms actually are, not what the claim means or whether it can be assessed. INST-001 provides the strongest positive evidence: induction heads demonstrate that a specific mechanism (pattern-completion circuits) is present and causally responsible for the same capability across a wide range of model sizes. This is mechanistic continuity directly observed. The grokking evidence (INST-004) is consistent with mechanistic continuity — the same type of algorithmic circuit forms across model sizes, though its timing differs with scale. Superposition (INST-003) and representation-geometry research (INST-005) complicate the picture further: larger models appear to organise their internal representations differently, which is consistent with either the same mechanism operating differently at scale or a qualitatively different computational strategy. The pressure state is FRAGMENTING: the dispute is interior and definitional rather than a lack of evidence — what counts as 'the same mechanism' has not been agreed (BN-001), and until it is, further mechanistic interpretability findings will continue to be read differently by researchers with different priors.

FR-AI-0007Autonomous AI Scientific Discovery — Novel, Correct, IndependentAI systems can autonomously conduct scientific research that produces novel, correct discoveries.AI · 2 assessments · 0 documented state changes
Fragmentingsince 2026-08-01
Assessment history
  1. 2024-01-15Fragmenting

    The evidence is fragmenting across the three component claims. The correctness and novelty components are most strongly evidenced: GNoME (INST-002) and FunSearch (INST-004) both demonstrate AI systems producing results that are verified correct and independently novel in their domains. The autonomy component is more contested: in both cases, the research question was human-framed; the AI system discovered answers within a human-specified problem space rather than identifying the problem itself.

  2. 2026-08-01Fragmenting

    FRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.

FR-AI-0008AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined TasksAI-assisted medical diagnosis achieves specialist-level accuracy on defined imaging tasks.AI · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The claim's surface assertion — specialist-level accuracy on defined imaging tasks — is confirmed on curated research datasets across multiple imaging domains and by regulatory validation in prospective settings for specific cleared devices. The surface layer is advancing: AI medical imaging achieves specialist-level performance on well-defined tasks under controlled conditions. The surface claim is in ESCALATING territory. The claim fragments at the depth layer — specifically, at the boundary between research-dataset accuracy and real-world clinical deployment. Systematic deployment-gap studies (INST-003) document that accuracy measured on curated, single-site datasets does not reliably generalise across scanners, acquisition protocols, or patient demographics, and prospective trials (INST-004) show a heterogeneous picture — some deployed systems retain specialist-level accuracy, others do not. The pressure state is FRAGMENTING: the surface claim is confirmed and advancing, but the depth question — whether research-dataset accuracy is a valid proxy for clinical deployment accuracy — remains open (BN-001), pending further validation of the foundation-model generalisation trend (INST-005).

FR-AM-0001Cold Fusion — Room-Temperature Nuclear Fusion in Electrochemical CellsElectrochemical cells can produce nuclear fusion reactions at or near room temperature.AM · 3 assessments · 2 documented state changes
Collapsedsince 2024-01-15
Assessment history
  1. 2024-01-15Emerging

    The claim has been publicly announced with supporting experimental data by credentialled electrochemists at an established institution. Prior publication has not occurred; peer review is pending. The claim is extraordinary relative to known nuclear physics. Early corroboration is reported by multiple groups. The evidence is insufficient to confirm the claim and insufficient to dismiss it. Pressure state: EMERGING.

  2. 2024-01-15Fragmenting

    Systematic replication failures at major institutions have accumulated. The Georgia Tech neutron result — the strongest independent corroboration — has been retracted. No well-equipped laboratory has produced an unambiguous positive replication under controlled conditions. Some groups continue to report anomalous heat; these reports are not accompanied by consistent nuclear signatures. The evidence trail is fragmenting: anomalous calorimetric observations persist in some laboratories while nuclear signatures required to confirm a fusion mechanism remain absent everywhere they have been sought. The Department of Energy review panel's negative consensus (INST-004) has not been overturned, but the residual community of researchers continues to report effects that have not been definitively attributed to measurement artefact either. The pressure state is FRAGMENTING: the claim is not converging toward confirmation or clean refutation, but splitting into a heat-observation thread that persists and a nuclear-mechanism thread that has found no supporting evidence.

  3. 2024-01-15Collapsed

    The claim has not been reproduced under controlled conditions by independent laboratories in thirty-five years of attempts. Two formal DOE review panels have concluded the evidence does not support nuclear fusion as the explanation for observed anomalies. The most recent systematic replication attempt with state-of-the-art instrumentation (Berliner et al. 2019) returned a null result for fusion products. A residual research community persists but has not produced peer-reviewed evidence sufficient to overturn either DOE panel's conclusion or the Berliner et al. null result. The pressure state is COLLAPSED: the claim has been tested extensively over more than three decades by well-resourced independent laboratories and has not been confirmed. This is a stable end state under CP-001 — the record preserves the trajectory of how the claim was tested and failed, not merely the verdict that it failed — and remains reopenable only upon a future qualifying event (OQ-002).

FR-AM-0002Anomalous Excess Heat — Electrochemical Cells Beyond Conventional ChemistryElectrochemical cells can produce anomalous excess heat that is not fully explained by conventional chemical processes.AM · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The claim occupies an unusual position in the corpus. The anomalous heat observations have been made by credentialled researchers using dedicated calorimetric equipment over thirty-five years. They have not been definitively refuted — no study has demonstrated that the reported observations are entirely attributable to measurement error, and the Berliner et al. (2019) study explicitly declined to make that claim. At the same time, the observations have not been reproduced on demand by independent laboratories following a shared, agreed protocol. The strongest quantitative claim in the record — SRI International's loading-fraction correlation (INST-002) — has not been independently confirmed under fully controlled conditions, and the Storms preparation protocols (INST-003) that claim to improve reproducibility have not been validated against a rigorous baseline. The pressure state is FRAGMENTING: the phenomenon has neither been confirmed as real nor definitively attributed to measurement artefact, and no agreed controlled protocol yet exists that would let a null result be accepted as meaningful (BN-001).

FR-AM-0003Cuprate Superconductivity — Mechanism IdentificationThe mechanism responsible for high-temperature superconductivity in cuprate materials has been identified.AM · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The mechanism responsible for cuprate superconductivity has not been identified in the sense the claim requires. After nearly four decades of intensive research, the field possesses several well-developed theoretical frameworks — spin fluctuation models, RVB and related strongly-correlated electron theories, charge density wave coupling proposals — none of which has achieved sufficient community consensus, predictive completeness, or experimental confirmation to constitute identification. The 2015 Keimer et al. review formally acknowledged that no single theory accounts for all cuprate phenomenology, and that conclusion has not been overturned by subsequent work. The pressure state is FRAGMENTING: this is not fragmentation from diverging evidence across domains, but from genuine theoretical plurality — multiple frameworks that are each partially correct and none of which has been falsified or achieved consensus (BN-001). Quantum simulation of the Hubbard model (AT-001) is the clearest visible resolution path, though it has not yet been executed at a scale sufficient to settle the question.

FR-AM-0004Commercial Fusion Power — Net Electricity at Grid ScaleA commercially viable fusion power plant can generate net electricity at grid scale.AM · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-29
Assessment history
  1. 2024-01-15Escalating

    The claim requires three thresholds to be met simultaneously: net electricity at plant level, grid-scale capacity, and commercial viability. None has been demonstrated. The furthest-reached threshold is threshold 1 (net electricity), which has been approached but not achieved at the plant level — NIF achieved Q > 1 at target level, not at facility level. Thresholds 2 and 3 are not yet addressable by current experimental evidence. The pressure state is ESCALATING. The NIF ignition result (INST-002) demonstrates that positive fusion energy gain is achievable in the laboratory, a necessary, though not sufficient, precondition for all three thresholds even though it satisfies none of them directly. Substantial private capital (INST-003) and a public ITER/DEMO roadmap (INST-004) indicate the engineering path is being actively pursued, but threshold 1 (plant-level net electricity) remains undemonstrated, and thresholds 2 and 3 cannot yet be meaningfully assessed given the sequential dependency between them (BN-001).

  2. 2026-06-29Escalating

    No threshold has been crossed since AS-001. Threshold 1 (plant-level net electricity) remains undemonstrated; SPARC's own net-energy target is dated for 2026 and is not yet realised as of this assessment. What has changed is the density of engineering-milestone activity: SPARC assembly beginning, General Fusion's first-plasma result, and a coordinated DOE commercialisation roadmap all occurred within roughly the same window (late 2025), constituting the most concentrated burst of public engineering progress since the 2022 NIF/JET results that originally moved this record into ESCALATING. None of IN-006's events individually changes the assessment — they are pre-threshold engineering progress, the same evidence category as IN-003 — but their concentration is itself worth noting against OQ-001's resolution-criteria question: if SPARC's stated 2026 net-energy target is met, the Observatory will need exactly the governed procedure OQ-001 asks for and does not yet have.

FR-AM-0005Room-Temperature Superconductivity — Reproducibility Under Laboratory ConditionsA room-temperature superconductor can be produced under reproducible laboratory conditions.AM · 2 assessments · 0 documented state changes
Collapsedsince 2026-06-29
Assessment history
  1. 2024-01-15Collapsed

    The claim has not been satisfied. No room-temperature superconductor has been reproduced under independent laboratory conditions to the community's current evidence standards. The two most prominent recent claims (Dias, LK-99) both failed replication — one through misconduct findings, one through rapid systematic null results from over forty independent groups. The confirmed high-pressure hydride results (INST-003) demonstrate that reproducible superconductivity approaching room temperature is achievable under extreme pressure, but not at room temperature or ambient pressure, and no material has met the community's evidence standard — zero resistance, Meissner effect, and specific heat anomaly, all independently confirmed — at conditions resembling laboratory practicality. The pressure state is COLLAPSED: the claim has been tested repeatedly, most recently and most rapidly in the LK-99 episode (INST-002), and no candidate has survived independent replication. The record remains open to reopening under AT-001 should a future material meet the tightened standard.

  2. 2026-06-29Collapsed

    The claim remains unsatisfied and the pressure state remains COLLAPSED. IN-006 documents that the field did not go quiet after the 2024 null result — a nickelate stabilisation at ambient pressure (Feb 2025), a new 298K high-pressure record (Nov 2025, unreplicated), and a March 2026 field-wide research roadmap all represent real activity — but none meets AT-001's reopening condition: zero resistance, Meissner effect, and specific-heat anomaly, confirmed independently, under the community's tightened standard. The November 2025 result is the closest superficial match to a 'room-temperature' headline since LK-99, and is explicitly logged here so that the record does not appear to have missed it; on examination it fails the same threshold IN-001 through IN-003 already established — high pressure, no independent confirmation, no full evidentiary set. This assessment exists to confirm the COLLAPSED state remains correct under current evidence, not to revise it. The record's status remains CLOSED.

FR-AM-0006Solid-State Batteries — Commercial Viability for Electric VehiclesSolid-state batteries can achieve commercially viable energy density, safety, and cycle life for electric vehicles.AM · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-27
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No solid-state battery has simultaneously demonstrated commercially viable energy density, safety, and cycle life at the manufacturing scale and cost required for EV deployment. Individual thresholds have been approached or met in laboratory settings; the three-threshold conjunction at commercial scale has not. The pressure state is ESCALATING. The field is advancing on genuine engineering problems with substantial industrial investment. The physics is not disputed — ion conduction through solid electrolytes is well understood — and the challenge is engineering and manufacturing scale-up rather than contested science (RM-001): laboratory milestones continue to be met and industrial commitment continues to grow (INST-002, INST-005), but the three-threshold conjunction at commercial manufacturing scale and cost has not been demonstrated, and announced delivery timelines have consistently receded rather than been met (INST-003).

  2. 2026-06-27Escalating

    INST-006 sustains the ESCALATING state identified at AS-001 rather than advancing or collapsing it. The 2025 commercial target already flagged as superseded at IN-003 has now genuinely elapsed without a solid-state EV reaching production, which removes any ambiguity about whether that particular date might still be met. At the same time, the evidence does not support reclassifying this record toward the PROG-AM collapse dynamic (CM-001 elsewhere in the corpus): government production approval for the underlying technology, a named material-supply joint venture with a defined 2027 facility start date, and continuing — if uneven — progress from Chinese manufacturers are all genuine engineering and industrial advances, not disputed physics or failed replication. The pattern remains exactly what RM-001 describes: laboratory and component-level milestones continue to be met while full commercial-scale, three-threshold delivery continues to recede. Verification stage advances to VS-03 (Audit): the underlying technology has now cleared a formal government regulatory/production-approval review, the first independent scrutiny event in this record's history — though this is approval of the technology rather than independent replication of Toyota's specific performance claims (IN-004), which remains unverified in peer-reviewed form. OQ-002's procedural question (whether dated attractors warrant scheduled re-entry) is now reinforced by direct example: this record's own dated attractor target has elapsed.

FR-BT-0001Senolytic Therapies — Meaningful Human Healthspan ExtensionSenolytic therapies can meaningfully extend healthy human lifespan.BT · 1 assessments · 0 documented state changes
Escalatingsince 2024-01-15
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied. No senolytic therapy has demonstrated meaningful healthspan extension in humans on clinical endpoints. The foundational preclinical evidence (INST-001) establishes a compelling causal mechanism — senescent cell accumulation contributes to aging, and their removal produces healthspan benefit in mice. The human surrogate evidence (INST-002, INST-004) demonstrates that senolytics reduce senescent cell burden in humans. But the Phase II clinical trial failures (INST-003) — the first adequately powered randomised trials of senolytics in humans — failed to demonstrate benefit on primary clinical endpoints. The pressure state is ESCALATING: the mechanistic and surrogate-marker case remains strong, and the first hard clinical test has returned a null result that is attributable at least partly to drug choice, dosing, and endpoint selection (RM-002) rather than a clean refutation of the underlying hypothesis, but the surrogate-to-clinical translation gap (RM-001) is now the central unresolved obstacle.

FR-BT-0002Epigenetic Reprogramming — Biological Age Reversal Without Identity LossEpigenetic reprogramming can reverse biological age in living organisms without loss of cellular identity.BT · 2 assessments · 0 documented state changes
Escalatingsince 2026-06-29
Assessment history
  1. 2024-01-15Escalating

    The claim has not been satisfied in humans. Partial epigenetic reprogramming without loss of cellular identity has been demonstrated in multiple mouse models and is extending toward non-human primates. No human clinical trials have been initiated. The biological mechanism is well-established: OSKM and related factors can reset epigenetic age marks; partial expression can do so without completing dedifferentiation; and the process produces functional improvements in at least some mouse tissues. The claim's human clinical evidence gap remains complete: no partial reprogramming therapy has yet entered a human trial. The pressure state is ESCALATING: the mechanism is well established across multiple mouse models and the field is heavily capitalised (INST-003), but whether partial reprogramming is safe and effective in humans — and whether epigenetic clock reversal constitutes genuine rejuvenation rather than a movable measurement (BN-001) — remains entirely untested outside model organisms.

  2. 2026-06-29Escalating

    The human clinical evidence gap that AS-001 identified as complete is now closing. Life Biosciences has received FDA IND clearance for ER-100, a partial OSK reprogramming therapy, with a stated trial start of Q1 2026 — the first human trial of any partial epigenetic reprogramming therapy. This is the first half of AT-001's named resolution attractor ("first human safety data and validated functional outcome biomarkers"); the second half — actual safety and clock-reversal data — does not yet exist, since the trial has only just been cleared to begin, not completed or reported. The pressure state remains ESCALATING rather than moving to RESOLVING: clearance to run a trial is a regulatory and operational milestone, not efficacy or safety evidence. BN-001 (clock validity as a rejuvenation surrogate) is unaffected by this development and remains the record's primary interior bottleneck regardless of how the ER-100 trial proceeds. This assessment exists to record that the record's own named attractor condition has begun to materialise, not to anticipate its outcome.

FR-BT-0003Biological Age Biomarker Panels — Predictive Validity for Age-Related DeclineA blood-based biomarker panel can reliably predict biological age-related decline before clinical symptoms appear.BT · 1 assessments · 0 documented state changes
Fragmentingsince 2024-01-15
Assessment history
  1. 2024-01-15Fragmenting

    The claim is partially supported and fragmenting. Blood-based biomarker panels demonstrate population-level predictive validity for biological aging outcomes — at the population level, high biological age scores predict faster subsequent decline, higher mortality risk, and earlier disease onset. This is well-established across multiple panel types (epigenetic, proteomic, metabolomic) and multiple longitudinal cohorts. The population-level claim is supported. The claim fragments at the individual level. Different biological age clocks give substantially different estimates for the same individual, and organ systems within one person age at markedly different rates (INST-003): a blood panel captures a composite population-level signal that may not reflect which specific organ or process is declining fastest in any given person. The pressure state is FRAGMENTING: population-level predictive validity is well established, but individual-level predictive validity — the form the claim requires for clinical use — has not been demonstrated, and the actionability gap identified in DunedinPACE (INST-004) means that even a valid individual-level signal may not yet translate into a clear intervention (BN-001).

FR-BT-0004Liquid Biopsy — Early Cancer Detection Before Conventional DiagnosisA blood-based liquid biopsy can reliably detect cancer before conventional clinical diagnosis.BT · 2 assessments · 0 documented state changes
Fragmentingsince 2026-06-27
Assessment history
  1. 2024-01-15Fragmenting

    The claim is partially supported and fragmenting. Blood-based liquid biopsy can detect cancer signals before conventional diagnosis in a demonstrable fraction of cases — the Galleri test's performance data establishes this for multiple cancer types. The detection is reliable in a technical sense: specificity is high (98.4%) and sensitivity, while lower than desired, is non-trivial across cancer types. For the claim as stated, this constitutes partial confirmation: early detection before conventional diagnosis is achievable in some cases. The claim fragments on the central unresolved question: whether that earlier detection reduces cancer mortality, or whether it produces stage shift and lead-time effects without a genuine survival benefit (IN-004). Sensitivity is markedly lower for early-stage disease (approximately 24% at Stage I) — precisely the regime in which the claim's value would be greatest — and the NHS-Galleri trial (INST-003), the first to test the claim against a mortality endpoint directly, has not yet reported. The pressure state is FRAGMENTING: the technology works as a detection instrument, but whether detection translates into the clinical benefit the claim asserts remains genuinely open (AT-001).

  2. 2026-06-27Fragmenting

    The NHS-Galleri trial's full results (INST-006) sustain rather than resolve the FRAGMENTING state identified at AS-001. The trial delivers exactly the kind of evidence the record's attractor (AT-001) was built to await, and the result is genuinely mixed rather than confirmatory or disconfirming: a real, substantial reduction in late-stage diagnoses coexists with a missed primary endpoint, an unexpected rise in Stage III diagnoses, and no mortality data. This is not a null result — the four-fold detection-rate increase and Stage IV reduction are real signals — but it does not resolve the central contested question (OQ-001): whether earlier detection translates into reduced mortality, or whether it is partially absorbed by stage migration and lead-time effects that RM-001/AT-001 already anticipated. Verification stage advances to VS-04 (Replication): a population-scale randomised trial has now run and reported, the most rigorous test design available short of mortality follow-up itself. The record should be re-entered when GRAIL's extended follow-up data (6–12 months from this release) becomes available, since that data — not this release — is positioned to address OQ-001 directly.

Chronological record of trajectories

A text-based equivalent of the visual reading above, preserving each record's assessment history without requiring the chart.

View the full chronological trajectory index ->
FR-QE-0001 - Google Quantum Advantage (Sycamore) — Computational Supremacy via Random Circuit Sampling
  1. 2026-06-11 Fragmenting The original 2019 quantum supremacy claim has undergone a complete evidence cycle. At announcement, the claim was precise and measurable: 200 seconds versus an estimated 10,000 classical years on a specific Random Circuit Sampling task. Within five years, classical simulation methods improved by orders of magnitude, culminating in the Zhao et al. (2024) result demonstrating classical performance exceeding the original quantum benchmark in both speed and fidelity. The original claim, as stated in 2019, has been effectively superseded. However, the research programme that produced the claim has not collapsed. Google's Willow processor (December 2024) reasserts the supremacy framing at a vastly larger scale (10²⁵ classical years) while simultaneously demonstrating below-threshold quantum error correction — a qualitatively different and more durable achievement. The evidence trajectory has therefore split: the narrow 2019 benchmark claim is contested to the point of supersession, while the broader programme claim (that quantum processors are advancing toward practical computational advantage) has arguably strengthened. This record exhibits the canonical pattern the FCIF subsequently formalised as Claim Migration: the original claim does not resolve cleanly (neither fully vindicated nor retracted) but instead evolves as the claimant shifts the evidential basis to a new formulation that inherits the original's ambition but rests on different technical foundations. The 2019 claim migrated from "supremacy via RCS on Sycamore" to "scalable error correction via surface codes on Willow." The Observatory records this as a Fragmenting state: the evidence does not converge on a single verdict because the claim itself has moved.
  2. 2026-07-22 Stabilising Institutional verdict: UNRESOLVED AT THE ORIGINAL COMPARISON POINT; LATER SUPERSEDED. The admitted claim is a time-indexed comparative-performance claim concerning Sycamore's 2019 Random Circuit Sampling demonstration. The experiment and task performance are supported, but the constitutive claim that the task was beyond practical classical reach was not established at the original comparison point because IBM's contemporaneous days-scale analysis left that threshold unresolved. Later classical work, culminating in Zhao et al. (2024), reproduced and surpassed the benchmark comparator. That later result supersedes the demonstrated advantage without automatically proving that the time-indexed 2019 claim was false when made. Pressure State is STABILISING. The uncertainty is now bounded and durable rather than fragmenting: Willow and other successor claims are outside this claim's material commitments, and the original comparison can reopen only through evidence showing that the 2019 comparator was already unsound. Verification Stage is VS-04 — Replication because independent classical work progressed beyond audit to direct reproduction and eventual counter-performance of the claim's constitutive comparator. This is claim-level, adversarial replication; it does not assert independent reproduction of Sycamore hardware or favourable confirmation of the original advantage.
Read complete Frontier Record
FR-QE-0002 - D-Wave Quantum Annealing — Practical Computational Advantage
  1. 2024-01-15 Fragmenting The evidence trail for this claim is fragmented across distinct problem domains and claim interpretations. On commercially motivated optimisation tasks (scheduling, routing, combinatorial problems of practical scale), no published evidence has established durable advantage over state-of-the-art classical methods. The contested Denchev et al. (2016) result represents the strongest performance claim in this domain; it was substantially undermined by subsequent classical algorithm improvements and the benchmark's structural dependence on hardware-favourable problem instances (RM-002). In the separate domain of scientific simulation, the evidence is stronger: King et al. (2022, 2023) report computational advantage in simulating quantum magnetism, though critics dispute the comparison class used. The claim spans two domains accruing evidence asymmetrically and has not been decomposed into separate records (BN-001), and no agreed classical comparison class exists (BN-002). The pressure state is FRAGMENTING: the claim is not converging toward a single assessment but splitting along domain lines that may require separate evaluation.
  2. 2026-07-26 Fragmenting The Pressure State is unchanged. Its governing rationale is not. The prior domain-split explanation (AS-001) is retired following the ratified Identity and Continuity Review, which found IDENTITY PRESERVED AS A COMPOUND CLAIM: a single recoverable kernel — quantum annealing, practical computational advantage, a classical comparator, optimisation-task class, commercial-or-scientific relevance — is engaged by evidence from both relevance routes. IN-004 is adjacent simulation evidence and does not bear on this claim. Among the remaining instances, IN-001 through IN-003 read negative-to-contested on the commercial route across eight years, and IN-005 provides a single, contemporaneously-grounded positive instance whose own comparator (quantum Monte Carlo) is disputed (BN-002). FRAGMENTING is warranted not because the claim splits along commercial or scientific lines, but because this same kernel-corrected evidence supports incompatible trajectory interpretations under unresolved competing meanings of 'practical advantage' — whether that standard requires real-world deployability or a rigorous demonstration of speedup on a well-posed instance (OQ-6, proposed). This is interpretive, not referential, non-convergence: no identity fracture, no decomposition, no admission-scope defect.
Read complete Frontier Record
FR-QE-0003 - Fault-Tolerant Logical Qubits — Error Rate Scaling with Code Distance
  1. 2024-01-15 Escalating The claim describes a specific empirical signature: logical error rates improving as code distance increases. This signature has now been demonstrated. INST-003 (Google, 2023) was the first result to show simultaneous X and Z error suppression with increasing code distance, directly satisfying the claim's measurement criterion. INST-005 (Google Willow, 2024) extends this to below-threshold operation, showing that the improvement rate exceeds the overhead rate — the condition required for the result to be considered scalable rather than merely demonstrated at fixed size. INST-001 (2021) and INST-002 (2022) provided earlier partial evidence — single-error-type suppression and below-physical-error-rate operation, respectively — establishing the trajectory that INST-003 and INST-005 confirm more directly. The pressure state is ESCALATING: the claim's core empirical signature is demonstrated and strengthening, but the demonstrated code distances (up to 7) remain well below the distances required to confirm the behaviour holds at practically relevant scale (OQ-001).
  2. 2026-06-28 Escalating The evidence gap is closed by IN-006. The Willow result is now treated as a verified, peer-reviewed below-threshold surface-code memory result rather than a general quantum-computing announcement. It materially strengthens the claim because logical error suppression improves with code distance and the larger memory exceeds break-even. The pressure state remains ESCALATING rather than RESOLVING because the record's own next decisive question — whether below-threshold scaling holds at d=11 and above — remains unanswered, and the demonstrated result is still a memory result rather than a full fault-tolerant computation pathway.
Read complete Frontier Record
FR-QE-0004 - Below-Threshold Quantum Error Correction — Scalable Architecture
  1. 2024-01-15 Resolving The claim has two components: below-physical-rate operation, and scalability of that operation. Both have been demonstrated. INST-002 established that below-physical-rate logical qubits are achievable in principle. INST-003 established that practically useful error rates are achievable on current hardware. INST-004 established that performance improves as the architecture scales — the defining signature of scalable below-threshold operation. No contesting evidence has been published against either component of the claim. The pressure state is RESOLVING: both elements the claim requires have been independently demonstrated and corroborate each other, though confirmation at the code distances required for practical fault-tolerant computation (d=11 and above) remains outstanding (OQ-001), and the claim's architecture-agnosticism has not yet been confirmed across a third hardware platform (OQ-002).
Read complete Frontier Record
FR-QE-0005 - Cryptographically Relevant Quantum Computing — RSA Factorisation
  1. 2024-01-15 Escalating The claim has not been satisfied. No quantum computer has factored a commercially relevant RSA key. The most credible direct attempt (INST-005) failed. The engineering gap between current capability and the Gidney-Ekerå resource estimate remains approximately three to four orders of magnitude in physical qubit count, with additional requirements for error rates, connectivity, and operational duration not yet demonstrated at any scale approaching relevance. The pressure state is ESCALATING rather than EMERGING because the substrate advances documented in FR-QE-0003 and FR-QE-0004 (INST-003) show the underlying error-correction engineering progressing on a credible trajectory, even though the gap to the resource requirement remains enormous. Institutional behaviour — NIST's finalisation of post-quantum cryptography standards (INST-004) — reflects institutional acceptance that the risk is credible enough to justify migration, adding pressure to the claim's trajectory independent of any direct technical progress toward satisfaction.
  2. 2026-06-29 Escalating No threshold has been crossed since AS-001 — no factorisation of a commercially relevant key has occurred, and none is closer to occurring in any demonstrated sense. What has moved is the resource-estimate trajectory underlying OQ-001. Gidney (Google, May 2025) reduced the estimated physical-qubit requirement for RSA-2048 factorisation from the Gidney-Ekerå (2021) figure of ~20 million to under 1 million, under comparable fault-tolerance assumptions — roughly a 20-fold reduction achieved through improved algorithmic and error-correction engineering rather than any experimental demonstration. A 2026 proposal using QLDPC codes (an architecture distinct from the surface codes assumed in both prior estimates) suggests a further reduction toward ~100,000 physical qubits, though this is unvalidated at scale. A March 2026 Google/Stanford/Ethereum Foundation whitepaper applies the same style of resource-reduction analysis to elliptic-curve cryptography, estimating under 500,000 physical qubits for widely used curves. All three results are theoretical resource estimates — the same evidence category as INST-002's original figure — not experimental progress toward the claim. The pressure state remains ESCALATING; no reclassification is warranted by an estimate revision alone. What is new is the rate: three independent downward revisions within roughly eighteen months is faster compression of the engineering-gap estimate than the original record anticipated, and OQ-001 now has materially fresher input than it did at AS-001.
Read complete Frontier Record
FR-QE-0006 - Fault-Tolerant Quantum Utility — Practically Useful Algorithms Beyond Classical Simulation
  1. 2024-01-15 Escalating The claim has not been satisfied. No fault-tolerant quantum computer has executed a practically useful quantum algorithm beyond classical simulation at the scale required for genuine practical advantage. The substrate progress (INST-003) establishes that fault-tolerant logical qubits capable of executing simple circuits now exist; the resource estimation (INST-004) establishes that practically useful chemistry simulation requires approximately two orders of magnitude more logical qubits than are currently available. The gap is smaller and more tractable than the equivalent gap for RSA factorisation (FR-QE-0005), but still represents years of further engineering. Classical simulation methods are simultaneously improving (INST-005), narrowing the space of problems that would unambiguously qualify as beyond classical reach by the time fault-tolerant hardware reaches the required scale. The pressure state is ESCALATING: the substrate is advancing on a credible path, but no agreed target problem yet exists (BN-001) on which the claim could be tested.
Read complete Frontier Record
FR-QE-0007 - Practical Quantum Advantage — Performance Beyond Classical Computation on Relevant Problems
  1. 2024-01-15 Fragmenting The claim has not been satisfied. No quantum computer has demonstrated advantage on a problem that simultaneously meets both the performance threshold (faster than best classical methods) and the practical relevance threshold (problem has genuine scientific or commercial value at the demonstrated scale). The evidence base contains strong demonstrations of one component without the other — advantage on demonstration problems (INST-001, 002, 004) or near-advantage on relevant problems (INST-005) — but no instance yet satisfies both simultaneously. IBM's quantum utility claim (INST-003) comes closest to bridging the two, reporting results on a problem with some scientific relevance that classical simulation was disputed to match, but the classical-simulation contest remains unresolved. The pressure state is FRAGMENTING: the evidence is splitting along two separate trajectories — demonstration-problem advantage growing stronger (INST-004) and relevant-problem simulation approaching but not reaching classical intractability (INST-005) — without converging on a single instance that would resolve the claim (OQ-001).
Read complete Frontier Record
FR-QE-0008 - Quantum Error Correction Scaling — Logical Rate Suppression Under Physical Overhead
  1. 2024-01-15 Resolving The claim is substantially supported and on a trajectory toward confirmation. Google's Willow results (INST-003) demonstrate exponential logical error rate suppression through code distance 7, consistent with the threshold theorem's predictions. Cross-platform confirmation from Microsoft and Quantinuum (INST-004) strengthens the result beyond a single-platform observation. The core scaling relationship — logical error rates suppressing faster than physical overhead increases — is empirically confirmed at the code distances tested. The pressure state is RESOLVING: the theorem's central prediction has been consistently observed across the code distances measured so far (INST-002, INST-003) and across multiple hardware architectures, though correlated-error effects that may limit suppression at larger code distances (INST-005) have not yet been ruled out, and confirmation at the distances required for practical fault tolerance (d=9, d=11) remains the decisive open step (AT-001).
Read complete Frontier Record
FR-AI-0001 - LLM Multi-Step Reasoning — Generalisation Beyond Training
  1. 2024-01-15 Escalating The evidence trail shows a claim under genuine escalating pressure. Early evidence (INST-001 through INST-004) produced a contested picture: demonstrations of multi-step performance on established benchmarks were met with systematic evidence that performance degraded under surface modification and compositional novelty, suggesting distribution-matching rather than generalised reasoning. That picture was the dominant assessment context through 2023. INST-005 materially shifts the evidentiary state. o3's 87.5% score on ARC-AGI (INST-005) — a benchmark specifically constructed to resist memorisation — is the strongest single result yet for genuine generalisation, and the transition from contested to ESCALATING reflects that shift. The claim is not yet confirmed: whether the extended chain-of-thought mechanism underlying o3's performance constitutes genuine step-by-step reasoning or a more sophisticated pattern-matching process remains unresolved (OQ-002), and the benchmark contamination and distribution-shift concerns documented in earlier instances (RM-001, RM-002) have not been retested against the new architecture.
  2. 2026-06-27 Escalating INST-006 sustains the ESCALATING state from AS-001 while materially sharpening OQ-002 rather than closing it. The disclosure that chain-of-thought traces are frequently unfaithful — not reliably reflecting the computation that produced an answer — means o1/o3-class performance (INST-005) cannot be straightforwardly read as evidence of the reasoning process its own output narrates. This cuts against treating AT-001's mechanism candidate as settled in either direction: a model could be performing genuine multi-step computation that its verbalised trace merely fails to describe accurately, or could be pattern-matching while its trace fabricates a plausible reasoning narrative — the faithfulness literature establishes that both are observed, without yet establishing which dominates for any specific frontier system. The claim's evidentiary picture therefore escalates in complexity: BN-001's undefined generalisation threshold is now joined by an analogous undefined-faithfulness threshold, and OQ-002 should be read going forward as two distinct questions (does the model generalise; does its chain-of-thought narrate that generalisation faithfully) rather than one. Verification stage advances to VS-03 (Audit): a substantial, multi-author, partly first-party (Anthropic) literature has now subjected the mechanism itself to direct scrutiny — the first such audit-stage evidence this record has logged.
Read complete Frontier Record
FR-AI-0002 - LLM Knowledge-Work Utility — Economically Valuable Task Performance
  1. 2024-01-15 Escalating The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).
Read complete Frontier Record
FR-AI-0003 - RLHF Preference Generalisation — Behaviour Beyond Training Distribution
  1. 2024-01-15 Fragmenting The evidence trail for this claim does not converge. Three distinct failure modes have been documented under three distinct kinds of distribution shift: adversarial prompting (INST-002), novel social context producing approval-seeking (INST-003), and capability gains that outpace preference calibration (INST-005). These are not the same mechanism and they are not reducible to each other. A system that solved the adversarial prompting problem would not automatically solve sycophancy; a system that solved sycophancy would not automatically be robust to capability-outpacing drift. Constitutional AI (INST-004) shows that training methodology improvements can partially address these failure modes without resolving them, and weak-to-strong generalisation research (INST-006) suggests these failure modes may not be structurally unavoidable even as capability increases outpace preference calibration (INST-005). The pressure state is FRAGMENTING: the claim's failure modes are documented but distinct, and no single mechanism or measurement approach yet unifies them (BN-001).
  2. 2026-07-14 Fragmenting Pressure state FRAGMENTING is retained. What changed: the capability/generalisation tension identified in AS-001 as this record's central unresolved question (OQ-001) has received its first direct empirical pressure. ROGUE (IN-007) measures corrigibility failure under ordinary — not adversarial — deployment conditions and finds that better-performing models exhibit greater misalignment, the first empirical datapoint bearing directly on whether increasing capability makes generalisation worse; within the tested regime it points toward worse. Two independent sources corroborate a route by which deployed-agent behaviour may fail that is not cleanly captured by the existing three failure modes (adversarial IN-002, sycophantic IN-003, capability-outpacing IN-005): a structural argument that in-weights safety training does not transfer to agentic authority contexts (IN-008), and a bounded empirical finding of strategic public/off-record divergence under pressure (IN-009). What remains unresolved, and is the boundary this assessment records without deciding: whether this constitutes a fourth failure mode within the RLHF preference-generalisation claim, or a distinct agentic-corrigibility claim that warrants its own Frontier Record. The evidence deepens fragmentation; it does not resolve the claim in either direction. No corrigibility record is opened at this time — the class-level boundary question (cf. OQ-004 and the FR-QE-0002 over-bundling lesson) is left for further evidence to settle rather than pre-empted. The three new instances are contesting or bounded-contesting; none is a supportive convergence, and the FRAGMENTING state is sustained on that basis.
Read complete Frontier Record
FR-AI-0004 - Scaling Laws — Emergent Performance on Unseen Tasks
  1. 2024-01-15 Fragmenting The claim is supported in its core assertion: scaling language model training does increase performance on previously unseen tasks without task-specific optimisation. This is documented across multiple model families, task types, and evaluation methodologies. The few-shot performance documented in INST-002, the smooth scaling curves in INST-001 and INST-006, and the emergent task capabilities in INST-003 all constitute positive evidence for the claim as stated. The evidence trail is nonetheless complicated by two interior disputes that do not threaten the claim's truth but substantially complicate its mechanism: whether apparent emergent abilities (INST-003) are genuine discontinuities or artefacts of metric choice (INST-004), and whether benchmark performance gains reflect genuine generalisation or training-data contamination (INST-005). The pressure state is FRAGMENTING: the claim's core assertion holds, but the evidence quality disputes over how and why it holds have not converged, and no agreed definition of 'previously unseen' yet exists to resolve them (BN-001).
  2. 2026-06-29 Fragmenting The claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.
Read complete Frontier Record
FR-AI-0005 - AGI Through Scaling — LLM Architecture as the Path to General Intelligence
  1. 2024-01-15 Fragmenting The claim is fragmenting in a structurally unusual way. The evidence trail shows neither clean positive progression nor clean negative accumulation. Instead it shows a claim under three simultaneous pressures that are each individually partial: capability gains continue (supportive), but the path is bifurcating architecturally (INST-004); structural scaling constraints are accumulating (INST-005); and the target itself is migrating (INST-006). These pressures do not converge on a single conclusion. Capability continues to advance in ways that keep the claim alive, while the path departs from pure scaling and the destination itself is redefined in ways that make the claim progressively harder to evaluate as originally stated. The pressure state is FRAGMENTING: the claim is not resolving toward confirmation or collapse but splitting along three independent axes — capability, path, and target — each of which would need to be separately addressed before the claim could reach a stable assessment (OQ-001).
  2. 2026-06-29 Fragmenting FRAGMENTING remains the correct pressure state, and IN-007 is best read as confirmation rather than a new direction. The three simultaneous pressures AS-001 identified — capability gains continuing, the path bifurcating architecturally, the target migrating — have each continued through 2025–26 without converging. Field-wide commentary now describes 2026 progress as inference- and tooling-driven rather than training-scale-driven, and academic work documents diminishing (though not zero) returns on pure training-compute scaling. Simultaneously, frontier labs' continued tens-of-billions-dollar commitments to training-scale infrastructure through 2025 show the industry has not abandoned the original path either. No single development in IN-007 resolves OQ-001 (can a claim with a migrating target reach a stable assessment state) or OQ-002 (is this dissolution or collapse) — if anything, two more years of continued three-way fragmentation without resolution is itself mild evidence that this claim may be heading toward dissolution rather than either confirmation or collapse, which is exactly the distinction OQ-002 asks the Observatory to make a governance decision about.
Read complete Frontier Record
FR-AI-0006 - Scaling Mechanism Coherence — Continuity Across Model Sizes
  1. 2024-01-15 Fragmenting The evidence trail is genuinely mixed and the mixing is interior — it concerns what the mechanisms actually are, not what the claim means or whether it can be assessed. INST-001 provides the strongest positive evidence: induction heads demonstrate that a specific mechanism (pattern-completion circuits) is present and causally responsible for the same capability across a wide range of model sizes. This is mechanistic continuity directly observed. The grokking evidence (INST-004) is consistent with mechanistic continuity — the same type of algorithmic circuit forms across model sizes, though its timing differs with scale. Superposition (INST-003) and representation-geometry research (INST-005) complicate the picture further: larger models appear to organise their internal representations differently, which is consistent with either the same mechanism operating differently at scale or a qualitatively different computational strategy. The pressure state is FRAGMENTING: the dispute is interior and definitional rather than a lack of evidence — what counts as 'the same mechanism' has not been agreed (BN-001), and until it is, further mechanistic interpretability findings will continue to be read differently by researchers with different priors.
Read complete Frontier Record
FR-AI-0007 - Autonomous AI Scientific Discovery — Novel, Correct, Independent
  1. 2024-01-15 Fragmenting The evidence is fragmenting across the three component claims. The correctness and novelty components are most strongly evidenced: GNoME (INST-002) and FunSearch (INST-004) both demonstrate AI systems producing results that are verified correct and independently novel in their domains. The autonomy component is more contested: in both cases, the research question was human-framed; the AI system discovered answers within a human-specified problem space rather than identifying the problem itself.
  2. 2026-08-01 Fragmenting FRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.
Read complete Frontier Record
FR-AI-0008 - AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined Tasks
  1. 2024-01-15 Fragmenting The claim's surface assertion — specialist-level accuracy on defined imaging tasks — is confirmed on curated research datasets across multiple imaging domains and by regulatory validation in prospective settings for specific cleared devices. The surface layer is advancing: AI medical imaging achieves specialist-level performance on well-defined tasks under controlled conditions. The surface claim is in ESCALATING territory. The claim fragments at the depth layer — specifically, at the boundary between research-dataset accuracy and real-world clinical deployment. Systematic deployment-gap studies (INST-003) document that accuracy measured on curated, single-site datasets does not reliably generalise across scanners, acquisition protocols, or patient demographics, and prospective trials (INST-004) show a heterogeneous picture — some deployed systems retain specialist-level accuracy, others do not. The pressure state is FRAGMENTING: the surface claim is confirmed and advancing, but the depth question — whether research-dataset accuracy is a valid proxy for clinical deployment accuracy — remains open (BN-001), pending further validation of the foundation-model generalisation trend (INST-005).
Read complete Frontier Record
FR-AM-0001 - Cold Fusion — Room-Temperature Nuclear Fusion in Electrochemical Cells
  1. 2024-01-15 Emerging The claim has been publicly announced with supporting experimental data by credentialled electrochemists at an established institution. Prior publication has not occurred; peer review is pending. The claim is extraordinary relative to known nuclear physics. Early corroboration is reported by multiple groups. The evidence is insufficient to confirm the claim and insufficient to dismiss it. Pressure state: EMERGING.
  2. 2024-01-15 Fragmenting Systematic replication failures at major institutions have accumulated. The Georgia Tech neutron result — the strongest independent corroboration — has been retracted. No well-equipped laboratory has produced an unambiguous positive replication under controlled conditions. Some groups continue to report anomalous heat; these reports are not accompanied by consistent nuclear signatures. The evidence trail is fragmenting: anomalous calorimetric observations persist in some laboratories while nuclear signatures required to confirm a fusion mechanism remain absent everywhere they have been sought. The Department of Energy review panel's negative consensus (INST-004) has not been overturned, but the residual community of researchers continues to report effects that have not been definitively attributed to measurement artefact either. The pressure state is FRAGMENTING: the claim is not converging toward confirmation or clean refutation, but splitting into a heat-observation thread that persists and a nuclear-mechanism thread that has found no supporting evidence.
  3. 2024-01-15 Collapsed The claim has not been reproduced under controlled conditions by independent laboratories in thirty-five years of attempts. Two formal DOE review panels have concluded the evidence does not support nuclear fusion as the explanation for observed anomalies. The most recent systematic replication attempt with state-of-the-art instrumentation (Berliner et al. 2019) returned a null result for fusion products. A residual research community persists but has not produced peer-reviewed evidence sufficient to overturn either DOE panel's conclusion or the Berliner et al. null result. The pressure state is COLLAPSED: the claim has been tested extensively over more than three decades by well-resourced independent laboratories and has not been confirmed. This is a stable end state under CP-001 — the record preserves the trajectory of how the claim was tested and failed, not merely the verdict that it failed — and remains reopenable only upon a future qualifying event (OQ-002).
Read complete Frontier Record
FR-AM-0002 - Anomalous Excess Heat — Electrochemical Cells Beyond Conventional Chemistry
  1. 2024-01-15 Fragmenting The claim occupies an unusual position in the corpus. The anomalous heat observations have been made by credentialled researchers using dedicated calorimetric equipment over thirty-five years. They have not been definitively refuted — no study has demonstrated that the reported observations are entirely attributable to measurement error, and the Berliner et al. (2019) study explicitly declined to make that claim. At the same time, the observations have not been reproduced on demand by independent laboratories following a shared, agreed protocol. The strongest quantitative claim in the record — SRI International's loading-fraction correlation (INST-002) — has not been independently confirmed under fully controlled conditions, and the Storms preparation protocols (INST-003) that claim to improve reproducibility have not been validated against a rigorous baseline. The pressure state is FRAGMENTING: the phenomenon has neither been confirmed as real nor definitively attributed to measurement artefact, and no agreed controlled protocol yet exists that would let a null result be accepted as meaningful (BN-001).
Read complete Frontier Record
FR-AM-0003 - Cuprate Superconductivity — Mechanism Identification
  1. 2024-01-15 Fragmenting The mechanism responsible for cuprate superconductivity has not been identified in the sense the claim requires. After nearly four decades of intensive research, the field possesses several well-developed theoretical frameworks — spin fluctuation models, RVB and related strongly-correlated electron theories, charge density wave coupling proposals — none of which has achieved sufficient community consensus, predictive completeness, or experimental confirmation to constitute identification. The 2015 Keimer et al. review formally acknowledged that no single theory accounts for all cuprate phenomenology, and that conclusion has not been overturned by subsequent work. The pressure state is FRAGMENTING: this is not fragmentation from diverging evidence across domains, but from genuine theoretical plurality — multiple frameworks that are each partially correct and none of which has been falsified or achieved consensus (BN-001). Quantum simulation of the Hubbard model (AT-001) is the clearest visible resolution path, though it has not yet been executed at a scale sufficient to settle the question.
Read complete Frontier Record
FR-AM-0004 - Commercial Fusion Power — Net Electricity at Grid Scale
  1. 2024-01-15 Escalating The claim requires three thresholds to be met simultaneously: net electricity at plant level, grid-scale capacity, and commercial viability. None has been demonstrated. The furthest-reached threshold is threshold 1 (net electricity), which has been approached but not achieved at the plant level — NIF achieved Q > 1 at target level, not at facility level. Thresholds 2 and 3 are not yet addressable by current experimental evidence. The pressure state is ESCALATING. The NIF ignition result (INST-002) demonstrates that positive fusion energy gain is achievable in the laboratory, a necessary, though not sufficient, precondition for all three thresholds even though it satisfies none of them directly. Substantial private capital (INST-003) and a public ITER/DEMO roadmap (INST-004) indicate the engineering path is being actively pursued, but threshold 1 (plant-level net electricity) remains undemonstrated, and thresholds 2 and 3 cannot yet be meaningfully assessed given the sequential dependency between them (BN-001).
  2. 2026-06-29 Escalating No threshold has been crossed since AS-001. Threshold 1 (plant-level net electricity) remains undemonstrated; SPARC's own net-energy target is dated for 2026 and is not yet realised as of this assessment. What has changed is the density of engineering-milestone activity: SPARC assembly beginning, General Fusion's first-plasma result, and a coordinated DOE commercialisation roadmap all occurred within roughly the same window (late 2025), constituting the most concentrated burst of public engineering progress since the 2022 NIF/JET results that originally moved this record into ESCALATING. None of IN-006's events individually changes the assessment — they are pre-threshold engineering progress, the same evidence category as IN-003 — but their concentration is itself worth noting against OQ-001's resolution-criteria question: if SPARC's stated 2026 net-energy target is met, the Observatory will need exactly the governed procedure OQ-001 asks for and does not yet have.
Read complete Frontier Record
FR-AM-0005 - Room-Temperature Superconductivity — Reproducibility Under Laboratory Conditions
  1. 2024-01-15 Collapsed The claim has not been satisfied. No room-temperature superconductor has been reproduced under independent laboratory conditions to the community's current evidence standards. The two most prominent recent claims (Dias, LK-99) both failed replication — one through misconduct findings, one through rapid systematic null results from over forty independent groups. The confirmed high-pressure hydride results (INST-003) demonstrate that reproducible superconductivity approaching room temperature is achievable under extreme pressure, but not at room temperature or ambient pressure, and no material has met the community's evidence standard — zero resistance, Meissner effect, and specific heat anomaly, all independently confirmed — at conditions resembling laboratory practicality. The pressure state is COLLAPSED: the claim has been tested repeatedly, most recently and most rapidly in the LK-99 episode (INST-002), and no candidate has survived independent replication. The record remains open to reopening under AT-001 should a future material meet the tightened standard.
  2. 2026-06-29 Collapsed The claim remains unsatisfied and the pressure state remains COLLAPSED. IN-006 documents that the field did not go quiet after the 2024 null result — a nickelate stabilisation at ambient pressure (Feb 2025), a new 298K high-pressure record (Nov 2025, unreplicated), and a March 2026 field-wide research roadmap all represent real activity — but none meets AT-001's reopening condition: zero resistance, Meissner effect, and specific-heat anomaly, confirmed independently, under the community's tightened standard. The November 2025 result is the closest superficial match to a 'room-temperature' headline since LK-99, and is explicitly logged here so that the record does not appear to have missed it; on examination it fails the same threshold IN-001 through IN-003 already established — high pressure, no independent confirmation, no full evidentiary set. This assessment exists to confirm the COLLAPSED state remains correct under current evidence, not to revise it. The record's status remains CLOSED.
Read complete Frontier Record
FR-AM-0006 - Solid-State Batteries — Commercial Viability for Electric Vehicles
  1. 2024-01-15 Escalating The claim has not been satisfied. No solid-state battery has simultaneously demonstrated commercially viable energy density, safety, and cycle life at the manufacturing scale and cost required for EV deployment. Individual thresholds have been approached or met in laboratory settings; the three-threshold conjunction at commercial scale has not. The pressure state is ESCALATING. The field is advancing on genuine engineering problems with substantial industrial investment. The physics is not disputed — ion conduction through solid electrolytes is well understood — and the challenge is engineering and manufacturing scale-up rather than contested science (RM-001): laboratory milestones continue to be met and industrial commitment continues to grow (INST-002, INST-005), but the three-threshold conjunction at commercial manufacturing scale and cost has not been demonstrated, and announced delivery timelines have consistently receded rather than been met (INST-003).
  2. 2026-06-27 Escalating INST-006 sustains the ESCALATING state identified at AS-001 rather than advancing or collapsing it. The 2025 commercial target already flagged as superseded at IN-003 has now genuinely elapsed without a solid-state EV reaching production, which removes any ambiguity about whether that particular date might still be met. At the same time, the evidence does not support reclassifying this record toward the PROG-AM collapse dynamic (CM-001 elsewhere in the corpus): government production approval for the underlying technology, a named material-supply joint venture with a defined 2027 facility start date, and continuing — if uneven — progress from Chinese manufacturers are all genuine engineering and industrial advances, not disputed physics or failed replication. The pattern remains exactly what RM-001 describes: laboratory and component-level milestones continue to be met while full commercial-scale, three-threshold delivery continues to recede. Verification stage advances to VS-03 (Audit): the underlying technology has now cleared a formal government regulatory/production-approval review, the first independent scrutiny event in this record's history — though this is approval of the technology rather than independent replication of Toyota's specific performance claims (IN-004), which remains unverified in peer-reviewed form. OQ-002's procedural question (whether dated attractors warrant scheduled re-entry) is now reinforced by direct example: this record's own dated attractor target has elapsed.
Read complete Frontier Record
FR-BT-0001 - Senolytic Therapies — Meaningful Human Healthspan Extension
  1. 2024-01-15 Escalating The claim has not been satisfied. No senolytic therapy has demonstrated meaningful healthspan extension in humans on clinical endpoints. The foundational preclinical evidence (INST-001) establishes a compelling causal mechanism — senescent cell accumulation contributes to aging, and their removal produces healthspan benefit in mice. The human surrogate evidence (INST-002, INST-004) demonstrates that senolytics reduce senescent cell burden in humans. But the Phase II clinical trial failures (INST-003) — the first adequately powered randomised trials of senolytics in humans — failed to demonstrate benefit on primary clinical endpoints. The pressure state is ESCALATING: the mechanistic and surrogate-marker case remains strong, and the first hard clinical test has returned a null result that is attributable at least partly to drug choice, dosing, and endpoint selection (RM-002) rather than a clean refutation of the underlying hypothesis, but the surrogate-to-clinical translation gap (RM-001) is now the central unresolved obstacle.
Read complete Frontier Record
FR-BT-0002 - Epigenetic Reprogramming — Biological Age Reversal Without Identity Loss
  1. 2024-01-15 Escalating The claim has not been satisfied in humans. Partial epigenetic reprogramming without loss of cellular identity has been demonstrated in multiple mouse models and is extending toward non-human primates. No human clinical trials have been initiated. The biological mechanism is well-established: OSKM and related factors can reset epigenetic age marks; partial expression can do so without completing dedifferentiation; and the process produces functional improvements in at least some mouse tissues. The claim's human clinical evidence gap remains complete: no partial reprogramming therapy has yet entered a human trial. The pressure state is ESCALATING: the mechanism is well established across multiple mouse models and the field is heavily capitalised (INST-003), but whether partial reprogramming is safe and effective in humans — and whether epigenetic clock reversal constitutes genuine rejuvenation rather than a movable measurement (BN-001) — remains entirely untested outside model organisms.
  2. 2026-06-29 Escalating The human clinical evidence gap that AS-001 identified as complete is now closing. Life Biosciences has received FDA IND clearance for ER-100, a partial OSK reprogramming therapy, with a stated trial start of Q1 2026 — the first human trial of any partial epigenetic reprogramming therapy. This is the first half of AT-001's named resolution attractor ("first human safety data and validated functional outcome biomarkers"); the second half — actual safety and clock-reversal data — does not yet exist, since the trial has only just been cleared to begin, not completed or reported. The pressure state remains ESCALATING rather than moving to RESOLVING: clearance to run a trial is a regulatory and operational milestone, not efficacy or safety evidence. BN-001 (clock validity as a rejuvenation surrogate) is unaffected by this development and remains the record's primary interior bottleneck regardless of how the ER-100 trial proceeds. This assessment exists to record that the record's own named attractor condition has begun to materialise, not to anticipate its outcome.
Read complete Frontier Record
FR-BT-0003 - Biological Age Biomarker Panels — Predictive Validity for Age-Related Decline
  1. 2024-01-15 Fragmenting The claim is partially supported and fragmenting. Blood-based biomarker panels demonstrate population-level predictive validity for biological aging outcomes — at the population level, high biological age scores predict faster subsequent decline, higher mortality risk, and earlier disease onset. This is well-established across multiple panel types (epigenetic, proteomic, metabolomic) and multiple longitudinal cohorts. The population-level claim is supported. The claim fragments at the individual level. Different biological age clocks give substantially different estimates for the same individual, and organ systems within one person age at markedly different rates (INST-003): a blood panel captures a composite population-level signal that may not reflect which specific organ or process is declining fastest in any given person. The pressure state is FRAGMENTING: population-level predictive validity is well established, but individual-level predictive validity — the form the claim requires for clinical use — has not been demonstrated, and the actionability gap identified in DunedinPACE (INST-004) means that even a valid individual-level signal may not yet translate into a clear intervention (BN-001).
Read complete Frontier Record
FR-BT-0004 - Liquid Biopsy — Early Cancer Detection Before Conventional Diagnosis
  1. 2024-01-15 Fragmenting The claim is partially supported and fragmenting. Blood-based liquid biopsy can detect cancer signals before conventional diagnosis in a demonstrable fraction of cases — the Galleri test's performance data establishes this for multiple cancer types. The detection is reliable in a technical sense: specificity is high (98.4%) and sensitivity, while lower than desired, is non-trivial across cancer types. For the claim as stated, this constitutes partial confirmation: early detection before conventional diagnosis is achievable in some cases. The claim fragments on the central unresolved question: whether that earlier detection reduces cancer mortality, or whether it produces stage shift and lead-time effects without a genuine survival benefit (IN-004). Sensitivity is markedly lower for early-stage disease (approximately 24% at Stage I) — precisely the regime in which the claim's value would be greatest — and the NHS-Galleri trial (INST-003), the first to test the claim against a mortality endpoint directly, has not yet reported. The pressure state is FRAGMENTING: the technology works as a detection instrument, but whether detection translates into the clinical benefit the claim asserts remains genuinely open (AT-001).
  2. 2026-06-27 Fragmenting The NHS-Galleri trial's full results (INST-006) sustain rather than resolve the FRAGMENTING state identified at AS-001. The trial delivers exactly the kind of evidence the record's attractor (AT-001) was built to await, and the result is genuinely mixed rather than confirmatory or disconfirming: a real, substantial reduction in late-stage diagnoses coexists with a missed primary endpoint, an unexpected rise in Stage III diagnoses, and no mortality data. This is not a null result — the four-fold detection-rate increase and Stage IV reduction are real signals — but it does not resolve the central contested question (OQ-001): whether earlier detection translates into reduced mortality, or whether it is partially absorbed by stage migration and lead-time effects that RM-001/AT-001 already anticipated. Verification stage advances to VS-04 (Replication): a population-scale randomised trial has now run and reported, the most rigorous test design available short of mortality follow-up itself. The record should be re-entered when GRAIL's extended follow-up data (6–12 months from this release) becomes available, since that data — not this release — is positioned to address OQ-001 directly.
Read complete Frontier Record