← ObservatoryThe RecordFR-AI-0004
PROG-AI
FR-AI-0004

Scaling Laws — Emergent Performance on Unseen Tasks

Scaling language model training increases performance on previously unseen tasks without task-specific optimisation.

FragmentingVS-03·since 2026-09-02
Assessment trajectory
Fragmentingstate held · last assessed 2026-09-02
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
VS-03
Audit
Current from 2024-01-15 — present
VS-04
Replication
VS-05
Operation
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateFragmentingVS-03
Why this state?Governed assessment correction following LPR-001-D04. AS-001 and AS-002 are preserved append-only as historical judgements; AS-003 supersedes only the source-fidelity claims identified by the provenance review. No new evidence instance was admitted through LPR-001.
Assessment summaryLPR-001-D04 corrects the source interpretation underlying parts of AS-001 and AS-002 without rewriting those historical assessments. The task-level core remains supported, principally by GPT-3 few-shot evaluations (IN-002) and the documented emergence literature (IN-003), but the evidence is narrower than AS-001 stated: Kaplan et al. (IN-001) establish predictable scaling of held-out language-model loss rather than unseen-task performance; Chinchilla (IN-006) establishes compute-optimal training and broad downstream gains rather than independently proving task novelty; and the contamination evidence (IN-005) establishes a serious measurement risk without showing that the record's cited benchmark gains were themselves contamination-driven. Schaeffer et al. (IN-004) narrow the emergence dispute to metric-dependent apparent discontinuities in particular evaluated settings, not a universal proof that scaling is continuous. Test-time compute (IN-007) is now anchored to primary evidence from Snell et al.; it is a genuine scaling mechanism but operates at inference time and therefore sits partly outside the claim's explicit training-scaling scope. The corrected evidence still does not converge on a clean account of what counts as a previously unseen task, how much benchmark evidence is contamination-free, or whether inference scaling belongs inside this claim. FRAGMENTING / VS-03 is therefore retained, with lower confidence in the stronger historical formulations but no basis for a state or stage change.
State entered2024-01-15
Last reaffirmed2026-09-02
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

Resistance MechanismRM-001

Training data contamination. For closed models, training data content is not fully disclosed. When a model performs well on a benchmark, the possibility that benchmark instances appeared in training data cannot be fully excluded. This mechanism creates a structural limitation on the evidential value of benchmark performance gains: positive results are always partially contestable on contamination grounds. The mechanism is more severe for the claim tracked here than for most, because the claim specifically asserts performance on previously unseen tasks — contamination directly challenges whether any given positive result satisfies the "previously unseen" criterion. The contamination resistance mechanism is not eliminable for closed models; it can only be reduced by purpose-built anti-contamination evaluation designs.

Resistance MechanismRM-002

Metric-dependence of apparent emergence. As documented by Schaeffer et al. (INST-004), apparent discontinuities in scaling curves are sensitive to metric choice. What appears as an emergent ability under one evaluation metric may appear as a smooth continuous improvement under another. This creates a resistance mechanism specific to the claim's language: "increases performance" can be measured in multiple ways, and the picture of what scaling achieves differs depending on measurement choice. The mechanism does not deny that scaling helps; it complicates how and when improvement is observed, which is a genuine interior difficulty for the claim's evidential base.

BottleneckBN-001

No agreed definition of "previously unseen." The claim requires that performance improvements occur on tasks the model has not been specifically trained on. But "previously unseen" is ambiguous across several dimensions: whether a task type appeared in training data, whether specific instances appeared, whether structural analogues appeared, and whether the model's in-context learning constitutes encountering the task during inference rather than training. Without an agreed operational definition, positive and contesting evidence cannot be cleanly compared. This is a bottleneck because it cannot be resolved by more experiments — it requires definitional agreement that currently does not exist. This bottleneck sits at the interior of the record, in measurement methodology, not at the claim identity or closure boundary.

AttractorAT-001

Open-weights models with disclosed training data. The contamination resistance mechanism (RM-001) and the "previously unseen" bottleneck (BN-001) both weaken substantially for models where training data is fully disclosed and verifiable. Open-weights models with documented training sets — LLaMA, Mistral, and similar — permit principled contamination analysis. As the open-weights ecosystem matures and evaluation methodology improves, the evidentiary base for the claim can become cleaner. The attractor is not a single experiment but a methodological development: a corpus of models where the "previously unseen" criterion can be operationally verified. This is an interior attractor — it resolves an evidence quality problem, not a boundary ambiguity.

Assessment History
2024-01-15
Initial assessment — Fragmenting
The claim is supported in its core assertion: scaling language model training does increase performance on previously unseen tasks without task-specific optimisation. This is documented across multiple model families, task types, and evaluation methodologies. The few-shot performance documented in INST-002, the smooth scaling curves in INST-001 and INST-006, and the emergent task capabilities in INST-003 all constitute positive evidence for the claim as stated. The evidence trail is nonetheless complicated by two interior disputes that do not threaten the claim's truth but substantially complicate its mechanism: whether apparent emergent abilities (INST-003) are genuine discontinuities or artefacts of metric choice (INST-004), and whether benchmark performance gains reflect genuine generalisation or training-data contamination (INST-005). The pressure state is FRAGMENTING: the claim's core assertion holds, but the evidence quality disputes over how and why it holds have not converged, and no agreed definition of 'previously unseen' yet exists to resolve them (BN-001).
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-06-29
Reassessed, no change — Fragmenting
The claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.
Sourced from: Medium, "The State of Large Language Models: Latest Updates & Trends (2025–2026)" (Feb 2026) for the inference/tooling consensus-shift framing; general field knowledge of o1 (Sept 2024), o3 (Dec 2024), and DeepSeek-R1 (Jan 2025) release timing and capability framing. The Medium source is a secondary roundup, not a primary research paper — adequate for establishing that a shift occurred, not for citing specific benchmark figures. Primary literature (e.g. the test-time-compute scaling papers referenced in FR-AI-0005's evidence trail) should be consulted before this assessment is extended with specific numbers.
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-09-02
Reassessed, no change — Fragmenting
LPR-001-D04 corrects the source interpretation underlying parts of AS-001 and AS-002 without rewriting those historical assessments. The task-level core remains supported, principally by GPT-3 few-shot evaluations (IN-002) and the documented emergence literature (IN-003), but the evidence is narrower than AS-001 stated: Kaplan et al. (IN-001) establish predictable scaling of held-out language-model loss rather than unseen-task performance; Chinchilla (IN-006) establishes compute-optimal training and broad downstream gains rather than independently proving task novelty; and the contamination evidence (IN-005) establishes a serious measurement risk without showing that the record's cited benchmark gains were themselves contamination-driven. Schaeffer et al. (IN-004) narrow the emergence dispute to metric-dependent apparent discontinuities in particular evaluated settings, not a universal proof that scaling is continuous. Test-time compute (IN-007) is now anchored to primary evidence from Snell et al.; it is a genuine scaling mechanism but operates at inference time and therefore sits partly outside the claim's explicit training-scaling scope. The corrected evidence still does not converge on a clean account of what counts as a previously unseen task, how much benchmark evidence is contamination-free, or whether inference scaling belongs inside this claim. FRAGMENTING / VS-03 is therefore retained, with lower confidence in the stronger historical formulations but no basis for a state or stage change.
Governed assessment correction following LPR-001-D04. AS-001 and AS-002 are preserved append-only as historical judgements; AS-003 supersedes only the source-fidelity claims identified by the provenance review. No new evidence instance was admitted through LPR-001.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

2017–19
Transformer scaling begins. GPT-1 and GPT-2 demonstrate that language model performance improves with scale. Task generalisation is observed informally but not systematically measured against scaling curves.
2020
Scaling laws quantified; GPT-3 demonstrates emergent few-shot performance. Kaplan et al. establish the power-law relationship. Brown et al. demonstrate that scale produces in-context task performance across diverse previously unseen tasks. The claim enters ESCALATING.
2022
Emergent abilities documented and immediately contested. Wei et al. document discontinuous emergence. Schaeffer et al. challenge the discontinuity interpretation. The mechanism debate opens as an interior question.
2022–23
Chinchilla refines scaling; contamination becomes systematic concern. Hoffmann et al. correct the parameter-data balance. Contamination literature matures. The claim enters FRAGMENTING as evidence quality disputes compound.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

Can an agreed operational definition of "previously unseen task" be established — one that resolves the contamination question and the structural-analogue question simultaneously? Without this, BN-001 cannot be closed regardless of experimental output. This is the primary interior bottleneck.

Raised 2024-01-15
OQ-002

Are apparent emergent abilities genuine discontinuities in capability, or continuous improvements made visible by nonlinear metrics? The Schaeffer et al. result is significant but contested. The answer matters for whether scaling law extrapolation can predict future task performance.

Raised 2024-01-15
OQ-003

PROG-AI now contains a substrate inversion: the foundational mechanism record (FR-AI-0004) is FRAGMENTING while downstream consequence records are ESCALATING. Does this inversion resolve as the scaling mechanism clarifies, or does the programme continue to build capability evidence on a contested mechanistic foundation?

Raised 2024-01-15
OQ-004

IN-007 introduces test-time/inference-time compute as a scaling axis the claim's original statement did not anticipate. Should the claim statement itself be revised to cover this — "scaling language model training and inference" — or does the Canonical Reality Principle's discipline (claims are stated as narrowly as the evidence warrants, scope changes go through the mutation log, not silent rewording) mean this is better tracked as an open question indefinitely, or spun into a related but distinct record? This is the same kind of structural question OQ-001 raises for "previously unseen," now recurring for "scaling."

Raised 2026-06-29
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0142026-09-06description_restoredLegacy ingestion cutoffs: mechanisms:RM-001, mechanisms:RM-002, mechanisms:BN-001, mechanisms:AT-001Source-restored complete descriptions
M-0132026-09-02assessment_issuedAS-002AS-003
M-0122026-09-02provenance_correctionLPR-001-D04 discrepancies_foundDISCREPANCIES-CORRECTED
M-0112026-09-02provenance_reviewLPR-001-D04
M-0102026-07-09description_reorderedDESCRIPTION-REORDERED
M-0092026-06-29open_question_raisedOQ-RAISED
M-0082026-06-29assessment_issuedAS-001AS-002
M-0072026-06-29instances_loggedINSTANCES-LOGGED
M-0062024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0052024-01-15null_boundary_condition_metNULL-BOUNDARY-CONDITION-MET
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
7 instances on recordShow sources ↓Hide ↑
IN-001Kaplan et al. — Neural scaling laws for language models1. Kaplan, J. et al. (2020), Scaling Laws for Neural Language Models, arXiv:2001.08361. · Abstract; cross-entropy-loss scaling with model size, dataset size and training computepartial
IN-002GPT-3 few-shot performance and emergent task generalisation1. Brown, T. B. et al. (2020), Language Models are Few-Shot Learners, NeurIPS 2020, arXiv:2005.14165. · Abstract; few-shot evaluation without gradient updates; translation, arithmetic and reasoning taskssupportive
IN-003Wei et al. — Emergent abilities of large language models1. Wei, J. et al. (2022), Emergent Abilities of Large Language Models, Transactions on Machine Learning Research, arXiv:2206.07682. · Definition of emergent abilities; examples and scaling discussionpartial
IN-004Schaeffer et al. — Are emergent abilities a mirage?1. Schaeffer, R., Miranda, B. and Koyejo, S. (2023), Are Emergent Abilities of Large Language Models a Mirage?, NeurIPS 2023, arXiv:2304.15004. · Abstract; metric-choice hypothesis; InstructGPT/GPT-3 and BIG-Bench analysescontesting
IN-005Benchmark contamination — evidence-quality constraint1. Golchin, S. and Surdeanu, M. (2023), Time Travel in LLMs: Tracing Data Contamination in Large Language Models, arXiv:2308.08493. · Abstract; contamination-detection method; GPT-4 findings for AG News, WNLI and XSumcontesting
IN-006Chinchilla — compute-optimal training1. Hoffmann, J. et al. (2022), Training Compute-Optimal Large Language Models, arXiv:2203.15556. · Abstract; compute-optimal scaling; Chinchilla downstream evaluationspartial
IN-007Test-time compute scaling — Snell et al.1. Snell, C. et al. (2024), Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, arXiv:2408.03314. · Abstract; compute-optimal test-time scaling; FLOPs-matched comparison with a 14x larger modelpartial