Large language models can perform multi-step reasoning that generalises beyond memorised training examples.
Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.
Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.
Benchmark contamination. For closed models, the content of training data is not publicly disclosed. Every benchmark performance result for a closed model carries an unresolvable uncertainty about whether the model encountered near-identical instances during training. This is a structural resistance mechanism: positive evidence on any established benchmark is always partially contestable on contamination grounds. The mechanism weakens as purpose-built anti-contamination benchmarks (ARC-AGI, FrontierMath) accumulate results, but is not eliminable for closed systems.
Distribution shift sensitivity. Performance on reasoning benchmarks consistently degrades under surface-level modifications that preserve logical structure. This is evidence that model performance tracks training distribution proximity rather than abstract reasoning rules. The mechanism has been robustly documented across multiple models and benchmark families (INST-002, INST-004). Its significance is contested — human performance also degrades under unfamiliar presentation — but it has not been refuted as a characterisation of current model behaviour.
Generalisation threshold is undefined. The claim requires that reasoning "generalises beyond memorised training examples" but no agreed standard specifies what degree of generalisation is sufficient, what comparison class defines memorisation, or what problem distribution constitutes "beyond training." Different researchers apply different implicit thresholds, producing genuine disagreement from shared evidence. This bottleneck differs from BN-001 in FR-QE-0002: it is not a domain scope ambiguity but a threshold ambiguity within a single domain. It is not resolvable by further evidence alone; it requires definitional consensus that does not currently exist.
Extended chain-of-thought as mechanism candidate. The o1/o3 architecture (INST-005) introduces explicit extended internal reasoning as a trained behaviour. If this mechanism produces genuine generalisation, it represents a qualitative shift in the kind of evidence available for this claim — not just better benchmark scores but a different computational process. Whether the mechanism produces true generalisation or a higher-order pattern-matching behaviour is the open research question. Recorded as an Attractor: it is drawing the evidence trail toward a potential resolution point, though the resolution is not yet achieved.
Historical narrative recorded for this claim. It does not override the current State Warrant.
Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.
Does ARC-AGI performance at 87.5% constitute sufficient evidence to confirm the claim, or does confirmation require performance across a wider distribution of anti-contamination benchmarks, including those not yet constructed?
Raised 2024-01-15Is the extended chain-of-thought mechanism in o1/o3 a qualitatively different computation from prior LLM inference, or a scaled version of the same pattern-matching behaviour documented in INST-002 and INST-004? This is the operative scientific question for the next assessment cycle.
Raised 2024-01-15Does the distribution shift sensitivity documented in INST-002 persist in o1/o3-class models? If it does, the claim must be scoped to specific problem classes rather than stated at the class level.
Raised 2024-01-15Can an agreed operational definition of "generalisation" be established that is acceptable to both the capability-affirming and capability-contesting research communities? Without this, BN-001 cannot be closed regardless of further evidence.
Raised 2024-01-15Is this a class-level claim or a system-level claim? The record tracks LLMs as a category, but the strongest positive evidence is from a specific architecture (o1/o3). If the generalisation mechanism is specific to the extended chain-of-thought training regime, the claim as stated may be too broad.
Raised 2024-01-15| Mutation | Date | Field | Prior value | Current value |
|---|---|---|---|---|
| M-011 | 2026-09-06 | description_restored | Legacy ingestion cutoffs: mechanisms:RM-001, mechanisms:RM-002, mechanisms:BN-001, mechanisms:AT-001, lineage:2017–20, lineage:2022, lineage:2023, lineage:2024 | Source-restored complete descriptions |
| M-010 | 2026-08-31 | provenance_review_completed | — | LPR-001-D02 |
| M-009 | 2026-08-29 | provenance_enriched | — | PROVENANCE-ENRICHED |
| M-008 | 2026-07-09 | description_reordered | — | DESCRIPTION-REORDERED |
| M-007 | 2026-06-27 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-006 | 2026-06-27 | instances_logged | — | INSTANCES-LOGGED |
| M-005 | 2024-01-15 | mechanisms_recorded | — | MECHANISMS-RECORDED |
| M-004 | 2024-01-15 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-003 | 2024-01-15 | scope_note_added | — | SCOPE-NOTE-ADDED |
| M-002 | 2024-01-15 | instances_logged | — | INSTANCES-LOGGED |
| M-001 | 2024-01-15 | record_created | — | RECORD-CREATED |