Capabilities that emerge through scaling language models are explained by the same underlying mechanism across model sizes.
Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.
Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.
Mechanistic interpretability does not yet scale to large models. The strongest positive evidence in this record (INST-001, induction heads) comes from mechanistic interpretability work conducted primarily on small to medium models (up to a few billion parameters). The techniques that identify and verify specific circuits — activation patching, attention head ablation, causal intervention — become computationally intractable at the scale of frontier models (100B+ parameters). Evidence about mechanisms in large models is therefore necessarily indirect: behavioural studies, representation geometry, and probing classifiers rather than direct circuit identification. The resistance mechanism is methodological: the tools that could directly test the claim are limited to the regime where the claim is most plausible and cannot currently reach the regime where it is most contested.
"Same mechanism" lacks an agreed operational definition. The claim requires that capabilities are explained by "the same underlying mechanism" across model sizes. Whether two mechanisms are "the same" depends on the level of abstraction at which identity is assessed. At the architectural level, all transformer models use the same mechanism (attention and MLP layers). At the circuit level, similar circuit types (induction heads, attention sinks) appear across scales. At the computational strategy level, large models may use their circuits differently enough to constitute mechanistic discontinuity. The claim cannot be evaluated — neither confirmed nor contested — without first agreeing on what level of mechanistic identity is required. This is the same lexical bottleneck type as FR-AI-0004 BN-001 and FR-AI-0005 BN-001, now appearing for the third time in PROG-AI.
Scalable mechanistic interpretability. The resolution path for this record is the development of mechanistic interpretability techniques that can operate at frontier model scale — identifying and verifying specific circuits in models with hundreds of billions of parameters. If such techniques are developed and applied to frontier models, they would either confirm that the same circuit types are causally responsible for capabilities at large scale (supporting the claim) or identify qualitatively different computational strategies (contesting it). The attractor is a methodological advance that would resolve RM-001. Several research groups (Anthropic, DeepMind, academic groups) are actively pursuing scalable interpretability; the timeline is unclear but the research direction is established.
Historical narrative recorded for this claim. It does not override the current State Warrant.
Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.
What level of mechanistic abstraction is required for "same mechanism" to be satisfied? Architectural (all transformers), circuit-type (induction heads appear at all scales), or computational strategy (circuits used similarly)? Until this is agreed, BN-001 cannot close regardless of experimental output.
Raised 2024-01-15PROG-AI now contains three records with lexical bottlenecks of the same type (FR-AI-0004, FR-AI-0005, FR-AI-0006). This is a programme-level pattern. Does it suggest that AI claims are particularly susceptible to lexical bottlenecks, or that the field is in an early stage where key terms have not yet been operationalised? Either interpretation has consequences for how the programme develops.
Raised 2024-01-15If scalable mechanistic interpretability is achieved (AT-001), would it resolve the claim? Or would the definitional bottleneck (BN-001) mean that even direct circuit evidence is interpreted differently by researchers with different priors about what "same mechanism" requires?
Raised 2024-01-15| Mutation | Date | Field | Prior value | Current value |
|---|---|---|---|---|
| M-016 | 2026-09-06 | description_restored | Legacy ingestion cutoffs: mechanisms:RM-001, mechanisms:BN-001, mechanisms:AT-001 | Source-restored complete descriptions |
| M-015 | 2026-09-04 | assessment_issued | AS-003 | AS-004 |
| M-014 | 2026-09-04 | instance_added | — | IN-007 |
| M-013 | 2026-09-04 | assessment_correction | AS-002 | AS-003 |
| M-012 | 2026-09-04 | provenance_correction | LPR-001-D06 discrepancies_found | LEGACY-INSTANCES-CORRECTED |
| M-011 | 2026-09-04 | provenance_review | — | LPR-001-D06 |
| M-010 | 2026-08-29 | provenance_enriched | IN-001 without structured provenance | IN-001 sources[] added |
| M-009 | 2026-08-29 | editorial_correction | IN-001 described causal responsibility across the full model-size range | IN-001 distinguishes causal small-model evidence from mainly correlational larger-model evidence |
| M-008 | 2026-08-29 | assessment_issued | AS-001 | AS-002 |
| M-007 | 2026-08-29 | instance_added | — | IN-006 |
| M-006 | 2024-01-15 | programme_panel_added | — | PROGRAMME-PANEL-ADDED |
| M-005 | 2024-01-15 | null_condition_met | — | NULL-CONDITION-MET |
| M-004 | 2024-01-15 | mechanisms_recorded | — | MECHANISMS-RECORDED |
| M-003 | 2024-01-15 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-002 | 2024-01-15 | instances_logged | — | INSTANCES-LOGGED |
| M-001 | 2024-01-15 | record_created | — | RECORD-CREATED |