Large language models can perform economically valuable knowledge-work tasks with limited human supervision.
Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.
Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.
Hallucination under low supervision. LLMs produce confident, fluent, and factually incorrect output at non-trivial rates. In supervised settings, human review catches and corrects errors before they produce harm. As supervision decreases, uncorrected errors propagate. The mechanism creates a hard constraint on the supervision level at which the claim holds: the claim is true under adequate supervision; the supervision threshold required varies by task stakes. This is not a reason to contest the claim but a reason to specify its domain of validity precisely.
Agentic error compounding. In single-turn assisted tasks, errors are local and recoverable. In multi-step agentic workflows, errors compound: a mistaken intermediate step propagates through subsequent steps, producing outputs that are difficult to audit and expensive to correct. The mechanism is evidenced by INST-006. It is structurally distinct from hallucination — it is not primarily about factual accuracy but about error propagation in sequential decision-making. As the frontier moves toward agentic deployment, this mechanism becomes the primary constraint on the claim's "limited supervision" condition.
Knowledge-work task scope is undefined. The claim spans a range from low-stakes, single-turn drafting assistance to high-stakes, multi-step autonomous professional work. The evidence is positive at the low end and negative or qualified at the high end. The bottleneck is not evidential — more evidence will not resolve it — it is definitional. A scoped version of the claim (e.g. restricted to single-turn assisted tasks in low-stakes domains) would already be confirmable. A broader version (including autonomous high-stakes professional work) would not. The Observatory records this without resolving it: the claim as stated spans both, and both are worth tracking.
Agentic capability threshold as resolution point. The primary open question for this record is whether LLMs can perform valuable knowledge-work autonomously over extended task sequences — the agentic frontier. Current evidence shows strong performance in supervised single-turn settings and weaker performance in autonomous multi-step settings. If agentic systems demonstrate reliable multi-step task completion in commercially deployed settings, the claim's upper scope boundary moves significantly. This is AT-001: the agentic capability frontier is drawing the evidence trail toward a potential resolution of BN-001. It is not yet reached; it is the attractor the record is currently moving toward.
Historical narrative recorded for this claim. It does not override the current State Warrant.
Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.
Where exactly is the supervision threshold boundary? The claim holds at supervised single-turn assistance and fails at autonomous multi-step high-stakes work. The boundary between these is not characterised. As agentic systems improve, does the boundary move, or is it a structural property of LLM architecture?
Raised 2024-01-15The productivity gains documented in INST-001, INST-002, and INST-005 are concentrated in tasks with clear outputs and measurable quality criteria (code, documents, customer service responses). Do the same gains hold for knowledge-work tasks with harder-to-measure outputs: strategic judgment, novel problem formulation, relationship management?
Raised 2024-01-15FR-AI-0001 (reasoning) and FR-AI-0002 (utility) are both ESCALATING. Is their co-escalation coincidental — both happen to be progressing simultaneously — or is there a causal relationship? If reasoning improvements (FR-AI-0001) drive utility improvements (FR-AI-0002), the two records are dependent. If utility advances independently of reasoning (through better prompting, deployment infrastructure, task design), they are not. This relationship is currently unobserved.
Raised 2024-01-15| Mutation | Date | Field | Prior value | Current value |
|---|---|---|---|---|
| M-011 | 2026-09-06 | description_restored | Legacy ingestion cutoffs: mechanisms:RM-001, mechanisms:RM-002, mechanisms:BN-001, mechanisms:AT-001, lineage:2020–22, lineage:2022–23, lineage:2023–24, lineage:2024 | Source-restored complete descriptions |
| M-010 | 2026-08-30 | provenance_review_completed | No governed provenance-review completion marker | LPR-001-D01 completed; discrepancies corrected |
| M-009 | 2026-08-30 | reference_corrected | IN-001 inferred comparable code quality; IN-002 used a secondary 37% speed expression; IN-004 misclassified Omiye et al. as a systematic review | IN-001 reports task success and the stated code-quality boundary; IN-002 uses the published 40% time reduction and 18% quality increase; IN-004 accurately describes Mata and the bounded Omiye et al. study; structured sources recorded |
| M-008 | 2026-08-17 | assessment_issued | AS-001 | AS-002 |
| M-007 | 2026-08-01 | instance_appended | IN-006 | IN-007 |
| M-006 | 2026-07-09 | description_reordered | — | DESCRIPTION-REORDERED |
| M-005 | 2024-01-15 | programme_panel_added | — | PROGRAMME-PANEL-ADDED |
| M-004 | 2024-01-15 | mechanisms_recorded | — | MECHANISMS-RECORDED |
| M-003 | 2024-01-15 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-002 | 2024-01-15 | instances_logged | — | INSTANCES-LOGGED |
| M-001 | 2024-01-15 | record_created | — | RECORD-CREATED |