← ObservatoryThe RecordFR-AI-0006
PROG-AI
FR-AI-0006

Scaling Mechanism Coherence — Continuity Across Model Sizes

Capabilities that emerge through scaling language models are explained by the same underlying mechanism across model sizes.

FragmentingVS-03·since 2026-09-04
Assessment trajectory
Fragmentingstate held · last assessed 2026-09-04
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
VS-03
Audit
Current from 2024-01-15 — present
VS-04
Replication
VS-05
Operation
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateFragmentingVS-03
Why this state?Bounded Record Review of Xu (2026), arXiv:2605.17084v1, originally surfaced as an LPR-001-D06 Record Review candidate. IN-007 is admitted as new evidence with primary provenance. It does not retroactively repair or replace unresolved legacy IN-005.
Assessment summaryNormal Record Review admits IN-007 as genuinely new cross-scale representation evidence. Xu's scale sweep supplies the type of direct small-versus-large comparison that the corrected legacy IN-003 through IN-005 lacked: predictive representation geometry follows different late-layer regimes in smaller and larger models. That result increases pressure on a strong continuity reading if mechanism identity is defined at the level of internal organisation. At the same time, the paper's own masking result cuts against treating the regime shift as proof of a wholly new mechanism: predictive structure remains recoverable beneath the dominant off-readout directions. IN-007 therefore deepens rather than resolves the record's central ambiguity. Together with IN-001 and IN-006, the evidence now contains meaningful support for both recurring structure and scale-associated specialisation, while BN-001 still prevents a stable answer to whether those observations count as the 'same underlying mechanism.' FRAGMENTING / VS-03 is retained. The evidential weight is moderated because IN-007 is a single-author preprint, introduces a new metric, and is not yet independently replicated.
State entered2024-01-15
Last reaffirmed2026-09-04
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

Resistance MechanismRM-001

Mechanistic interpretability does not yet scale to large models. The strongest positive evidence in this record (INST-001, induction heads) comes from mechanistic interpretability work conducted primarily on small to medium models (up to a few billion parameters). The techniques that identify and verify specific circuits — activation patching, attention head ablation, causal intervention — become computationally intractable at the scale of frontier models (100B+ parameters). Evidence about mechanisms in large models is therefore necessarily indirect: behavioural studies, representation geometry, and probing classifiers rather than direct circuit identification. The resistance mechanism is methodological: the tools that could directly test the claim are limited to the regime where the claim is most plausible and cannot currently reach the regime where it is most contested.

BottleneckBN-001

"Same mechanism" lacks an agreed operational definition. The claim requires that capabilities are explained by "the same underlying mechanism" across model sizes. Whether two mechanisms are "the same" depends on the level of abstraction at which identity is assessed. At the architectural level, all transformer models use the same mechanism (attention and MLP layers). At the circuit level, similar circuit types (induction heads, attention sinks) appear across scales. At the computational strategy level, large models may use their circuits differently enough to constitute mechanistic discontinuity. The claim cannot be evaluated — neither confirmed nor contested — without first agreeing on what level of mechanistic identity is required. This is the same lexical bottleneck type as FR-AI-0004 BN-001 and FR-AI-0005 BN-001, now appearing for the third time in PROG-AI.

AttractorAT-001

Scalable mechanistic interpretability. The resolution path for this record is the development of mechanistic interpretability techniques that can operate at frontier model scale — identifying and verifying specific circuits in models with hundreds of billions of parameters. If such techniques are developed and applied to frontier models, they would either confirm that the same circuit types are causally responsible for capabilities at large scale (supporting the claim) or identify qualitatively different computational strategies (contesting it). The attractor is a methodological advance that would resolve RM-001. Several research groups (Anthropic, DeepMind, academic groups) are actively pursuing scalable interpretability; the timeline is unclear but the research direction is established.

Assessment History
2024-01-15
Initial assessment — Fragmenting
The evidence trail is genuinely mixed and the mixing is interior — it concerns what the mechanisms actually are, not what the claim means or whether it can be assessed. INST-001 provides the strongest positive evidence: induction heads demonstrate that a specific mechanism (pattern-completion circuits) is present and causally responsible for the same capability across a wide range of model sizes. This is mechanistic continuity directly observed. The grokking evidence (INST-004) is consistent with mechanistic continuity — the same type of algorithmic circuit forms across model sizes, though its timing differs with scale. Superposition (INST-003) and representation-geometry research (INST-005) complicate the picture further: larger models appear to organise their internal representations differently, which is consistent with either the same mechanism operating differently at scale or a qualitatively different computational strategy. The pressure state is FRAGMENTING: the dispute is interior and definitional rather than a lack of evidence — what counts as 'the same mechanism' has not been agreed (BN-001), and until it is, further mechanistic interpretability findings will continue to be read differently by researchers with different priors.
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-08-29
Reassessed, no change — Fragmenting
IN-006 adds direct mechanistic evidence from abstract reasoning that specialized symbolic-processing circuitry is substantially associated with capable larger models and is weak or absent in smaller models that do not perform the task. This increases pressure on a simple cross-scale continuity reading, while not resolving the claim: the result is task-specific, and BN-001 remains decisive because whether an emergent specialized circuit counts as a new mechanism depends on the level of abstraction used for mechanism identity. FRAGMENTING is therefore retained. VS-03 is retained provisionally because the new evidence uses causal mediation and ablation-style mechanistic scrutiny; PA-005 does not reopen the historical stage classification beyond the evidence reviewed here.
PA-005 provenance-in-review replication trial. New evidence provenance captured at admission. Attempted opportunistic enrichment of IN-001 exposed a material wording/provenance discrepancy and was stopped for separate bounded correction review.
2026-09-04
Reassessed, no change — Fragmenting
LPR-001-D06 narrows the historical evidence underlying AS-001 without rewriting that assessment. Olsson et al. (IN-001) remain the strongest continuity evidence, but causal support is strongest in small attention-only models and becomes mainly correlational in larger models; the record therefore cannot describe cross-scale causal continuity as directly established. Wei and Michaud (IN-002) address behavioural emergence and a proposed quantized scaling model without determining mechanism identity across model sizes. Elhage et al. (IN-003) establish superposition in toy networks, and the grokking literature (IN-004) reverse-engineers gradual circuit formation in small transformers; neither supplies the cross-scale comparisons AS-001 previously inferred. The legacy representation-geometry bundle in IN-005 is withdrawn from the current evidential basis because its provenance could not be reconstructed. IN-006 remains direct cross-model evidence that specialized symbolic circuitry becomes more evident in capable larger models, increasing pressure on a simple continuity reading. FRAGMENTING / VS-03 is retained on this narrower basis: some continuity evidence exists, some task-specific evidence points toward scale-associated mechanistic specialization, and the governing definition of 'same mechanism' remains unresolved.
Governed assessment correction following LPR-001-D06. AS-001 and AS-002 remain visible append-only as historical judgements. No new evidence instance was admitted through LPR-001; the separately identified 2026 representation-geometry preprint remains a normal Record Review candidate.
2026-09-04
Reassessed, no change — Fragmenting
Normal Record Review admits IN-007 as genuinely new cross-scale representation evidence. Xu's scale sweep supplies the type of direct small-versus-large comparison that the corrected legacy IN-003 through IN-005 lacked: predictive representation geometry follows different late-layer regimes in smaller and larger models. That result increases pressure on a strong continuity reading if mechanism identity is defined at the level of internal organisation. At the same time, the paper's own masking result cuts against treating the regime shift as proof of a wholly new mechanism: predictive structure remains recoverable beneath the dominant off-readout directions. IN-007 therefore deepens rather than resolves the record's central ambiguity. Together with IN-001 and IN-006, the evidence now contains meaningful support for both recurring structure and scale-associated specialisation, while BN-001 still prevents a stable answer to whether those observations count as the 'same underlying mechanism.' FRAGMENTING / VS-03 is retained. The evidential weight is moderated because IN-007 is a single-author preprint, introduces a new metric, and is not yet independently replicated.
Bounded Record Review of Xu (2026), arXiv:2605.17084v1, originally surfaced as an LPR-001-D06 Record Review candidate. IN-007 is admitted as new evidence with primary provenance. It does not retroactively repair or replace unresolved legacy IN-005.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

2020–21
Scaling laws assume mechanistic continuity implicitly. Kaplan et al. scaling laws treat capability as a smooth function of scale, implicitly assuming continuous underlying mechanisms. The mechanistic question is not asked.
2022
Mechanistic interpretability makes parts of the continuity question empirically tractable. Olsson et al. identify induction heads with strong causal evidence in small attention-only models and mainly correlational evidence in larger models. Elhage et al. provide a toy-model account of superposition, relevant to representational mechanism but not a cross-scale language-model comparison.
2022–23
Emergent abilities and mechanistic explanations separate. Wei et al. document behavioural emergence with scale; Michaud et al. propose quantized skill acquisition as one explanation; grokking work shows apparently sudden behavioural transitions can arise from gradual circuit formation in small transformers. None of these results alone determines whether the same mechanism persists across model sizes.
2024–25
The legacy 2024 representation-geometry attribution cannot be confidently reconstructed and is withdrawn from the current evidential basis. In 2025, Yang et al. provide direct cross-model evidence that specialized symbolic mechanisms are associated with capable larger models, sharpening rather than resolving the continuity question.
2026
Xu introduces Subspace PGA and reports a scale-dependent regime in predictive representation geometry across seven Pythia models, with cross-family checks. Smaller models lose late-layer predictive alignment while larger models preserve it, yet the underlying predictive structure can be recovered after removing dominant off-readout directions. The result supplies direct cross-scale representation evidence while preserving the mechanism-identity ambiguity.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

What level of mechanistic abstraction is required for "same mechanism" to be satisfied? Architectural (all transformers), circuit-type (induction heads appear at all scales), or computational strategy (circuits used similarly)? Until this is agreed, BN-001 cannot close regardless of experimental output.

Raised 2024-01-15
OQ-002

PROG-AI now contains three records with lexical bottlenecks of the same type (FR-AI-0004, FR-AI-0005, FR-AI-0006). This is a programme-level pattern. Does it suggest that AI claims are particularly susceptible to lexical bottlenecks, or that the field is in an early stage where key terms have not yet been operationalised? Either interpretation has consequences for how the programme develops.

Raised 2024-01-15
OQ-003

If scalable mechanistic interpretability is achieved (AT-001), would it resolve the claim? Or would the definitional bottleneck (BN-001) mean that even direct circuit evidence is interpreted differently by researchers with different priors about what "same mechanism" requires?

Raised 2024-01-15
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0162026-09-06description_restoredLegacy ingestion cutoffs: mechanisms:RM-001, mechanisms:BN-001, mechanisms:AT-001Source-restored complete descriptions
M-0152026-09-04assessment_issuedAS-003AS-004
M-0142026-09-04instance_addedIN-007
M-0132026-09-04assessment_correctionAS-002AS-003
M-0122026-09-04provenance_correctionLPR-001-D06 discrepancies_foundLEGACY-INSTANCES-CORRECTED
M-0112026-09-04provenance_reviewLPR-001-D06
M-0102026-08-29provenance_enrichedIN-001 without structured provenanceIN-001 sources[] added
M-0092026-08-29editorial_correctionIN-001 described causal responsibility across the full model-size rangeIN-001 distinguishes causal small-model evidence from mainly correlational larger-model evidence
M-0082026-08-29assessment_issuedAS-001AS-002
M-0072026-08-29instance_addedIN-006
M-0062024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0052024-01-15null_condition_metNULL-CONDITION-MET
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
7 instances on recordShow sources ↓Hide ↑
IN-001Anthropic mechanistic interpretability — induction heads and in-context learning1. Olsson, C. et al. (2022), In-context Learning and Induction Heads, Transformer Circuits Thread. · Summary of Evidence for Sub-Claims; Arguments 1–6; Model Analysis Table2. Olsson, C. et al. (2022), In-context Learning and Induction Heads, arXiv:2209.11895. · Abstractsupportive
IN-002Emergent abilities and the quantization hypothesis — mechanism identity remains underdetermined1. Wei, J. et al. (2022), Emergent Abilities of Large Language Models, Transactions on Machine Learning Research, arXiv:2206.07682. · Definition of emergent abilities; scaling discussion2. Michaud, E. J. et al. (2023), The Quantization Model of Neural Scaling, NeurIPS 2023, arXiv:2303.13506. · Abstract; Quantization Hypothesis; toy-data validation and language-model decompositionpartial
IN-003Toy models of superposition — representational compression mechanism1. Elhage, N. et al. (2022), Toy Models of Superposition, Transformer Circuits Thread, arXiv:2209.10652. · Abstract; toy-model superposition and polysemanticity resultspartial
IN-004Grokking — gradual circuit formation behind delayed generalisation1. Power, A. et al. (2022), Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets, arXiv:2201.02177. · Abstract; delayed generalisation after overfitting2. Nanda, N. et al. (2023), Progress Measures for Grokking via Mechanistic Interpretability, ICLR 2023, arXiv:2301.05217. · Abstract; reverse-engineered modular-addition circuit; memorisation, circuit formation and cleanup phasespartial
IN-005Legacy representation-geometry attribution — provenance unresolvedpartial
IN-006Emergent symbolic mechanisms for abstract reasoning — cross-scale circuit evidence1. Yang, Y. et al. (2025), Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models, Proceedings of Machine Learning Research 267, ICML 2025. · Abstract and model-scale analysescontesting
IN-007Scale-dependent predictive representation geometry — Xu preprint1. Xu, W. (2026), Scale Determines Whether Language Models Organize Representation Geometry for Prediction, arXiv:2605.17084v1. · Abstract; §§4–5; seven-model Pythia scale sweep; cross-architecture validation; limitationspartial