Scaling language model training increases performance on previously unseen tasks without task-specific optimisation.
Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.
Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.
Training data contamination. For closed models, training data content is not fully disclosed. When a model performs well on a benchmark, the possibility that benchmark instances appeared in training data cannot be fully excluded. This mechanism creates a structural limitation on the evidential value of benchmark performance gains: positive results are always partially contestable on contamination grounds. The mechanism is more severe for the claim tracked here than for most, because the claim specifically asserts performance on previously unseen tasks — contamination directly challenges whether any given positive result satisfies the "previously unseen" criterion. The contamination resistance mechanism is not eliminable for closed models; it can only be reduced by purpose-built anti-contamination evaluation designs.
Metric-dependence of apparent emergence. As documented by Schaeffer et al. (INST-004), apparent discontinuities in scaling curves are sensitive to metric choice. What appears as an emergent ability under one evaluation metric may appear as a smooth continuous improvement under another. This creates a resistance mechanism specific to the claim's language: "increases performance" can be measured in multiple ways, and the picture of what scaling achieves differs depending on measurement choice. The mechanism does not deny that scaling helps; it complicates how and when improvement is observed, which is a genuine interior difficulty for the claim's evidential base.
No agreed definition of "previously unseen." The claim requires that performance improvements occur on tasks the model has not been specifically trained on. But "previously unseen" is ambiguous across several dimensions: whether a task type appeared in training data, whether specific instances appeared, whether structural analogues appeared, and whether the model's in-context learning constitutes encountering the task during inference rather than training. Without an agreed operational definition, positive and contesting evidence cannot be cleanly compared. This is a bottleneck because it cannot be resolved by more experiments — it requires definitional agreement that currently does not exist. This bottleneck sits at the interior of the record, in measurement methodology, not at the claim identity or closure boundary.
Open-weights models with disclosed training data. The contamination resistance mechanism (RM-001) and the "previously unseen" bottleneck (BN-001) both weaken substantially for models where training data is fully disclosed and verifiable. Open-weights models with documented training sets — LLaMA, Mistral, and similar — permit principled contamination analysis. As the open-weights ecosystem matures and evaluation methodology improves, the evidentiary base for the claim can become cleaner. The attractor is not a single experiment but a methodological development: a corpus of models where the "previously unseen" criterion can be operationally verified. This is an interior attractor — it resolves an evidence quality problem, not a boundary ambiguity.
Historical narrative recorded for this claim. It does not override the current State Warrant.
Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.
Can an agreed operational definition of "previously unseen task" be established — one that resolves the contamination question and the structural-analogue question simultaneously? Without this, BN-001 cannot be closed regardless of experimental output. This is the primary interior bottleneck.
Raised 2024-01-15Are apparent emergent abilities genuine discontinuities in capability, or continuous improvements made visible by nonlinear metrics? The Schaeffer et al. result is significant but contested. The answer matters for whether scaling law extrapolation can predict future task performance.
Raised 2024-01-15PROG-AI now contains a substrate inversion: the foundational mechanism record (FR-AI-0004) is FRAGMENTING while downstream consequence records are ESCALATING. Does this inversion resolve as the scaling mechanism clarifies, or does the programme continue to build capability evidence on a contested mechanistic foundation?
Raised 2024-01-15IN-007 introduces test-time/inference-time compute as a scaling axis the claim's original statement did not anticipate. Should the claim statement itself be revised to cover this — "scaling language model training and inference" — or does the Canonical Reality Principle's discipline (claims are stated as narrowly as the evidence warrants, scope changes go through the mutation log, not silent rewording) mean this is better tracked as an open question indefinitely, or spun into a related but distinct record? This is the same kind of structural question OQ-001 raises for "previously unseen," now recurring for "scaling."
Raised 2026-06-29| Mutation | Date | Field | Prior value | Current value |
|---|---|---|---|---|
| M-014 | 2026-09-06 | description_restored | Legacy ingestion cutoffs: mechanisms:RM-001, mechanisms:RM-002, mechanisms:BN-001, mechanisms:AT-001 | Source-restored complete descriptions |
| M-013 | 2026-09-02 | assessment_issued | AS-002 | AS-003 |
| M-012 | 2026-09-02 | provenance_correction | LPR-001-D04 discrepancies_found | DISCREPANCIES-CORRECTED |
| M-011 | 2026-09-02 | provenance_review | — | LPR-001-D04 |
| M-010 | 2026-07-09 | description_reordered | — | DESCRIPTION-REORDERED |
| M-009 | 2026-06-29 | open_question_raised | — | OQ-RAISED |
| M-008 | 2026-06-29 | assessment_issued | AS-001 | AS-002 |
| M-007 | 2026-06-29 | instances_logged | — | INSTANCES-LOGGED |
| M-006 | 2024-01-15 | programme_panel_added | — | PROGRAMME-PANEL-ADDED |
| M-005 | 2024-01-15 | null_boundary_condition_met | — | NULL-BOUNDARY-CONDITION-MET |
| M-004 | 2024-01-15 | mechanisms_recorded | — | MECHANISMS-RECORDED |
| M-003 | 2024-01-15 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-002 | 2024-01-15 | instances_logged | — | INSTANCES-LOGGED |
| M-001 | 2024-01-15 | record_created | — | RECORD-CREATED |