← ObservatoryThe RecordFR-AI-0005
PROG-AI
FR-AI-0005

AGI Through Scaling — LLM Architecture as the Path to General Intelligence

Artificial General Intelligence will be achieved through scaling current large-language-model architectures.

FragmentingVS-03·since 2026-09-03
Assessment trajectory
Fragmentingstate held · last assessed 2026-09-03
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
VS-03
Audit
Current from 2024-01-15 — present
VS-04
Replication
VS-05
Operation
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateFragmentingVS-03
Why this state?Governed assessment correction following LPR-001-D05. AS-001 and AS-002 remain visible as historical judgements. AS-003 supersedes the source-fidelity claims identified by the provenance review; no new evidence instance was admitted through LPR-001. Downstream mechanism, lineage, and open-question wording that independently relies on the withdrawn IN-006 target-migration premise requires a separate bounded consistency repair and is not silently changed here.
Assessment summaryLPR-001-D05 materially narrows the evidence underlying AS-001 and AS-002 without rewriting those historical assessments. The corrected record no longer supports their three-axis formulation of capability/path/target fragmentation: IN-006 does not document a 2024–25 migration of OpenAI's AGI definition, and the specific target-migration event is withdrawn. The evidence that remains is still non-convergent. GPT-3 and GPT-4 show substantial capability gains from large-scale language-model development (IN-001, IN-002), while Sutskever's later remarks, OpenAI o1, the finite-human-data analysis, and test-time-compute research indicate that the practical scaling recipe is changing and broadening beyond simply increasing pre-training scale (IN-003, IN-004, IN-005, IN-007). None of those results establishes that current LLM architectures will reach AGI, and none establishes a hard scaling ceiling or abandonment of LLM-based approaches. FRAGMENTING / VS-03 is therefore retained, but on a narrower basis: continuing capability gains coexist with unresolved path-definition and scaling-regime changes. The previous target-migration rationale should no longer be treated as evidential support for the current assessment.
State entered2024-01-15
Last reaffirmed2026-09-03
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

Resistance MechanismRM-001

Capability-generality gap. LLM scaling produces measurable improvements on benchmarks and economically valuable tasks but has not demonstrated the open-ended generalisation, physical understanding, causal reasoning, or sustained autonomy that earlier AGI framings required. The gap between benchmark performance and general intelligence has not closed with scale — it has become better characterised. Each capability advance reveals new gaps rather than eliminating the structural distance between current LLM behaviour and AGI as originally conceived. This mechanism is not about scaling hitting a ceiling; it is about the scaling path producing something systematically different from the target it was supposed to reach.

BottleneckBN-001

AGI definition is not stable. The claim requires that AGI be achieved, but AGI is not a stable object. Different researchers, institutions, and time periods use different definitions. The target has migrated at least once during this record's evidence trail (INST-006). A moving target creates a bottleneck that is not resolvable by experimental evidence: any positive result can be reframed as not yet AGI; any definitional narrowing can make existing systems qualify. The claim cannot transition to a stable assessment state — RESOLVING, COLLAPSED, or confirmed — until the target is fixed. This is a lexical bottleneck of the same type as BN-001 in FR-AI-0004, but more severe: the key term is not merely undefined, it is actively contested and institutionally managed.

BottleneckBN-002

Path and destination are disaggregating. The claim asserts a path to a destination. As the path evolves (toward hybrid architectures, test-time compute, embodiment, or novel approaches) while the destination also migrates, the original claim becomes progressively harder to evaluate. The path has bifurcated from pure LLM scaling; the destination has narrowed from general intelligence to economic performance. If both continue to move, the claim may become unevaluable not because evidence is absent but because the claim object no longer has a stable referent. This is a structural bottleneck specific to path prediction claims — a new bottleneck type the corpus has not previously named.

AttractorAT-001

Demonstrated AGI under a stable definition. The only clear resolution path for this record is a demonstration of AGI under a definition that the research community accepts as stable and meaningful — not the migrated economic performance definition, but a definition that captures the original intent of the claim. If such a demonstration occurred, the question of whether the path was pure LLM scaling or something that evolved from it would become secondary. The attractor is not a specific experiment but a definitional stabilisation followed by a capability demonstration. Whether either component is achievable on a near-term timeline is unclear.

Assessment History
2024-01-15
Initial assessment — Fragmenting
The claim is fragmenting in a structurally unusual way. The evidence trail shows neither clean positive progression nor clean negative accumulation. Instead it shows a claim under three simultaneous pressures that are each individually partial: capability gains continue (supportive), but the path is bifurcating architecturally (INST-004); structural scaling constraints are accumulating (INST-005); and the target itself is migrating (INST-006). These pressures do not converge on a single conclusion. Capability continues to advance in ways that keep the claim alive, while the path departs from pure scaling and the destination itself is redefined in ways that make the claim progressively harder to evaluate as originally stated. The pressure state is FRAGMENTING: the claim is not resolving toward confirmation or collapse but splitting along three independent axes — capability, path, and target — each of which would need to be separately addressed before the claim could reach a stable assessment (OQ-001).
Verification Stage: VS-03 preserved — historically unverified.
2026-06-29
Reassessed, no change — Fragmenting
FRAGMENTING remains the correct pressure state, and IN-007 is best read as confirmation rather than a new direction. The three simultaneous pressures AS-001 identified — capability gains continuing, the path bifurcating architecturally, the target migrating — have each continued through 2025–26 without converging. Field-wide commentary now describes 2026 progress as inference- and tooling-driven rather than training-scale-driven, and academic work documents diminishing (though not zero) returns on pure training-compute scaling. Simultaneously, frontier labs' continued tens-of-billions-dollar commitments to training-scale infrastructure through 2025 show the industry has not abandoned the original path either. No single development in IN-007 resolves OQ-001 (can a claim with a migrating target reach a stable assessment state) or OQ-002 (is this dissolution or collapse) — if anything, two more years of continued three-way fragmentation without resolution is itself mild evidence that this claim may be heading toward dissolution rather than either confirmation or collapse, which is exactly the distinction OQ-002 asks the Observatory to make a governance decision about.
Sourced from: Medium, "The State of Large Language Models: Latest Updates & Trends (2025–2026)" (Feb 2026); aimultiple.com summary of 2026 RL post-training scaling-laws research describing a "latent saturation trend"; Metaintro coverage of continued frontier-lab capital expenditure commitments through 2025 (Dec 2025). All three are secondary roundups rather than primary papers — adequate for establishing the shape of the 2025–26 debate, not for citing specific benchmark or expenditure figures as precise.
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-09-03
Reassessed, no change — Fragmenting
LPR-001-D05 materially narrows the evidence underlying AS-001 and AS-002 without rewriting those historical assessments. The corrected record no longer supports their three-axis formulation of capability/path/target fragmentation: IN-006 does not document a 2024–25 migration of OpenAI's AGI definition, and the specific target-migration event is withdrawn. The evidence that remains is still non-convergent. GPT-3 and GPT-4 show substantial capability gains from large-scale language-model development (IN-001, IN-002), while Sutskever's later remarks, OpenAI o1, the finite-human-data analysis, and test-time-compute research indicate that the practical scaling recipe is changing and broadening beyond simply increasing pre-training scale (IN-003, IN-004, IN-005, IN-007). None of those results establishes that current LLM architectures will reach AGI, and none establishes a hard scaling ceiling or abandonment of LLM-based approaches. FRAGMENTING / VS-03 is therefore retained, but on a narrower basis: continuing capability gains coexist with unresolved path-definition and scaling-regime changes. The previous target-migration rationale should no longer be treated as evidential support for the current assessment.
Governed assessment correction following LPR-001-D05. AS-001 and AS-002 remain visible as historical judgements. AS-003 supersedes the source-fidelity claims identified by the provenance review; no new evidence instance was admitted through LPR-001. Downstream mechanism, lineage, and open-question wording that independently relies on the withdrawn IN-006 target-migration premise requires a separate bounded consistency repair and is not silently changed here.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

2017–19
Scaling hypothesis implicit. GPT-1 and GPT-2 establish that scale improves language model performance. The AGI-through-scaling claim is implicit in the research programme but not yet explicitly stated as a path prediction.
2020–22
Path claim made explicit. GPT-3 and subsequent commentary formalise the claim. Leading researchers publicly assert that scaling LLMs is the path to AGI. The claim enters ESCALATING.
2023
First internal fracture. OpenAI board crisis and Sutskever's eventual departure signal that the originator community is beginning to disaggregate path from destination. The claim enters FRAGMENTING.
2024
Path bifurcation and definition migration. o1/o3 architectures depart from pure LLM scaling. OpenAI narrows the AGI definition to economic task performance. The claim's two components — path and destination — both migrate simultaneously.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

Can a claim whose target term is actively migrating reach any stable assessment state — RESOLVING, COLLAPSED, or confirmed — or does target migration structurally prevent closure? This is the collapse-criteria question this record generates. It is not the same as the resolution-criteria question in GQ-001, but it is related.

Raised 2024-01-15
OQ-002

Is claim dissolution — a claim becoming unevaluable because its referents have migrated — the same governance problem as claim collapse? Cold fusion collapsed through a defined external event. This claim may dissolve through definitional drift. If these are different objects, the Observatory may need to distinguish them.

Raised 2024-01-15
OQ-003

FR-AI-0005 is the first path prediction claim in the corpus. Path predictions age differently from capability claims: they can be overtaken by events, rendered moot by alternative paths succeeding, or abandoned by their proponents without formal falsification. Should path prediction claims be treated as a distinct record class, or does the current schema handle them adequately?

Raised 2024-01-15
OQ-004

Two years of continued three-way fragmentation (capability/path/target) without convergence is itself a data point. At what point — if any — should sustained non-convergence be treated as evidence toward dissolution (OQ-002's distinction) rather than simply more fragmentation? This record has no stated threshold for when 'still fragmenting' becomes 'has dissolved,' and OQ-002 already flagged that the Observatory lacks a governed answer to this question generally, not just for this record.

Raised 2026-06-29
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0122026-09-06description_restoredLegacy ingestion cutoffs: mechanisms:RM-001, mechanisms:BN-001, mechanisms:BN-002, mechanisms:AT-001Source-restored complete descriptions
M-0112026-09-03assessment_correctionAS-002AS-003
M-0102026-09-03provenance_correctionLPR-001-D05 discrepancies_foundLEGACY-INSTANCES-CORRECTED
M-0092026-09-03provenance_reviewLPR-001-D05
M-0082026-06-29open_question_raisedOQ-RAISED
M-0072026-06-29assessment_issuedAS-001AS-002
M-0062026-06-29instances_loggedINSTANCES-LOGGED
M-0052024-01-15null_condition_failedNULL-CONDITION-FAILED
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
7 instances on recordShow sources ↓Hide ↑
IN-001GPT-3 — scaling produces broad few-shot capability gains1. Brown, T. B. et al. (2020), Language Models are Few-Shot Learners, NeurIPS 2020, arXiv:2005.14165. · Abstract; few-shot and zero-shot evaluations without gradient updatespartial
IN-002GPT-4 — large capability gains with stated real-world limitations1. OpenAI (2023), GPT-4 Technical Report, arXiv:2303.08774. · Abstract; benchmark performance; real-world limitations; predictable scaling discussionpartial
IN-003Sutskever post-OpenAI position — scaling formula expected to change1. Reuters (14 Dec 2024), AI with reasoning power will be less predictable, Ilya Sutskever says. · NeurIPS 2024 remarks on pre-training limits, finite data and future reasoning systemspartial
IN-004OpenAI o1 — reasoning performance scales with train-time and test-time compute1. OpenAI (12 Sep 2024), Learning to reason with LLMs. · Training approach; chain of thought; scaling with reinforcement-learning and test-time compute; evaluationspartial
IN-005Human-generated training-data constraint1. Villalobos, P. et al. (2022), Will we run out of data? Limits of LLM scaling based on human-generated data, arXiv:2211.04325. · Abstract; public human-text stock forecast; 2026–2032 range; mitigation discussioncontesting
IN-006OpenAI AGI definition — longstanding economically valuable work formulation1. OpenAI, OpenAI Charter. · Mission statement defining AGI as highly autonomous systems outperforming humans at most economically valuable workpartial
IN-007Test-time compute — a second compute-scaling axis1. Snell, C. et al. (2024), Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, arXiv:2408.03314. · Abstract; compute-optimal test-time scaling; FLOPs-matched comparison with a 14x larger modelpartial