← ObservatoryThe RecordFR-AI-0002
PROG-AI
FR-AI-0002

LLM Knowledge-Work Utility — Economically Valuable Task Performance

Large language models can perform economically valuable knowledge-work tasks with limited human supervision.

EscalatingVS-02·since 2026-08-17
Assessment trajectory
Escalatingstate held · last assessed 2026-08-17
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
Current from 2024-01-15 — present
VS-03
Audit
VS-04
Replication
VS-05
Operation
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateEscalatingVS-02
Why this state?Issued during OHR-2026-09 catch-up review to close the unassessed-evidence gap created by IN-007. The reassessment preserves the distinction between bounded human-reviewed utility and extended low-supervision agentic work.
Assessment summaryThe claim remains ESCALATING and bounded by the supervision threshold identified in AS-001. IN-007 adds a stronger operational test at that boundary: on live-system root-cause analysis, general-purpose coding agents achieve low accuracy on realistic and hard tasks, hallucinate causes, and often miss concurrent incidents. This is meaningful contesting evidence against reliable delegation of extended, high-stakes multi-step knowledge work under low supervision. It does not overturn the claim's supported core because the record is not a universal claim of autonomous competence: controlled studies and deployed settings continue to show economically valuable performance where humans review outputs and the task is bounded. IN-007 therefore narrows confidence at the agentic frontier rather than reversing the underlying utility finding. Pressure State remains ESCALATING and Verification Stage remains VS-02; OQ-001 remains the decisive unresolved boundary.
State entered2024-01-15
Last reaffirmed2026-08-17
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

Resistance MechanismRM-001

Hallucination under low supervision. LLMs produce confident, fluent, and factually incorrect output at non-trivial rates. In supervised settings, human review catches and corrects errors before they produce harm. As supervision decreases, uncorrected errors propagate. The mechanism creates a hard constraint on the supervision level at which the claim holds: the claim is true under adequate supervision; the supervision threshold required varies by task stakes. This is not a reason to contest the claim but a reason to specify its domain of validity precisely.

Resistance MechanismRM-002

Agentic error compounding. In single-turn assisted tasks, errors are local and recoverable. In multi-step agentic workflows, errors compound: a mistaken intermediate step propagates through subsequent steps, producing outputs that are difficult to audit and expensive to correct. The mechanism is evidenced by INST-006. It is structurally distinct from hallucination — it is not primarily about factual accuracy but about error propagation in sequential decision-making. As the frontier moves toward agentic deployment, this mechanism becomes the primary constraint on the claim's "limited supervision" condition.

BottleneckBN-001

Knowledge-work task scope is undefined. The claim spans a range from low-stakes, single-turn drafting assistance to high-stakes, multi-step autonomous professional work. The evidence is positive at the low end and negative or qualified at the high end. The bottleneck is not evidential — more evidence will not resolve it — it is definitional. A scoped version of the claim (e.g. restricted to single-turn assisted tasks in low-stakes domains) would already be confirmable. A broader version (including autonomous high-stakes professional work) would not. The Observatory records this without resolving it: the claim as stated spans both, and both are worth tracking.

AttractorAT-001

Agentic capability threshold as resolution point. The primary open question for this record is whether LLMs can perform valuable knowledge-work autonomously over extended task sequences — the agentic frontier. Current evidence shows strong performance in supervised single-turn settings and weaker performance in autonomous multi-step settings. If agentic systems demonstrate reliable multi-step task completion in commercially deployed settings, the claim's upper scope boundary moves significantly. This is AT-001: the agentic capability frontier is drawing the evidence trail toward a potential resolution of BN-001. It is not yet reached; it is the attractor the record is currently moving toward.

Assessment History
2024-01-15
Initial assessment — Escalating
The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).
Verification Stage: VS-02 preserved — historically unverified.
2026-08-17
Reassessed, no change — Escalating
The claim remains ESCALATING and bounded by the supervision threshold identified in AS-001. IN-007 adds a stronger operational test at that boundary: on live-system root-cause analysis, general-purpose coding agents achieve low accuracy on realistic and hard tasks, hallucinate causes, and often miss concurrent incidents. This is meaningful contesting evidence against reliable delegation of extended, high-stakes multi-step knowledge work under low supervision. It does not overturn the claim's supported core because the record is not a universal claim of autonomous competence: controlled studies and deployed settings continue to show economically valuable performance where humans review outputs and the task is bounded. IN-007 therefore narrows confidence at the agentic frontier rather than reversing the underlying utility finding. Pressure State remains ESCALATING and Verification Stage remains VS-02; OQ-001 remains the decisive unresolved boundary.
Issued during OHR-2026-09 catch-up review to close the unassessed-evidence gap created by IN-007. The reassessment preserves the distinction between bounded human-reviewed utility and extended low-supervision agentic work.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

2020–22
GPT-3 and early commercial deployment. OpenAI API opens GPT-3 for commercial access. Early adopters deploy in copywriting, summarisation, and customer service contexts. Results are mixed; the capability exists but hallucination rates and inconsistency limit reliable deployment. The utility claim is in an EMERGING state: existence demonstrated, economic value unconfirmed at scale.
2022–23
ChatGPT release and mass adoption. ChatGPT reaches 100 million users in two months. Spontaneous mass deployment across knowledge-work contexts — drafting, coding, analysis, research — produces an enormous but methodologically uncontrolled evidence trail. Anecdotal reports of value are ubiquitous; systematic measurement is absent. The deployment scale constitutes a form of evidence but is not sufficient for the Observatory to issue a formal assessment.
2023
Controlled studies and commercial deployments. The Peng et al. and Noy and Zhang studies provide the first rigorous experimental evidence. Simultaneously, major professional services firms (law, consulting, finance) begin commercial deployment. The claim transitions to ESCALATING.
2023–24
Hallucination incidents and agentic frontier. High-profile hallucination failures in legal and medical contexts establish the supervision threshold as a genuine constraint. Agentic deployment attempts reveal the compounding error problem as a distinct resistance mechanism. The claim's boundary conditions become observable.
2024
Natural experiment at deployment scale. Brynjolfsson et al. publish the first large-scale natural experiment confirming productivity gains in deployed commercial settings. The claim's evidential base shifts from experimental to observational — from controlled trials to real economic outcomes in operating companies.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

Where exactly is the supervision threshold boundary? The claim holds at supervised single-turn assistance and fails at autonomous multi-step high-stakes work. The boundary between these is not characterised. As agentic systems improve, does the boundary move, or is it a structural property of LLM architecture?

Raised 2024-01-15
OQ-002

The productivity gains documented in INST-001, INST-002, and INST-005 are concentrated in tasks with clear outputs and measurable quality criteria (code, documents, customer service responses). Do the same gains hold for knowledge-work tasks with harder-to-measure outputs: strategic judgment, novel problem formulation, relationship management?

Raised 2024-01-15
OQ-003

FR-AI-0001 (reasoning) and FR-AI-0002 (utility) are both ESCALATING. Is their co-escalation coincidental — both happen to be progressing simultaneously — or is there a causal relationship? If reasoning improvements (FR-AI-0001) drive utility improvements (FR-AI-0002), the two records are dependent. If utility advances independently of reasoning (through better prompting, deployment infrastructure, task design), they are not. This relationship is currently unobserved.

Raised 2024-01-15
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0112026-09-06description_restoredLegacy ingestion cutoffs: mechanisms:RM-001, mechanisms:RM-002, mechanisms:BN-001, mechanisms:AT-001, lineage:2020–22, lineage:2022–23, lineage:2023–24, lineage:2024Source-restored complete descriptions
M-0102026-08-30provenance_review_completedNo governed provenance-review completion markerLPR-001-D01 completed; discrepancies corrected
M-0092026-08-30reference_correctedIN-001 inferred comparable code quality; IN-002 used a secondary 37% speed expression; IN-004 misclassified Omiye et al. as a systematic reviewIN-001 reports task success and the stated code-quality boundary; IN-002 uses the published 40% time reduction and 18% quality increase; IN-004 accurately describes Mata and the bounded Omiye et al. study; structured sources recorded
M-0082026-08-17assessment_issuedAS-001AS-002
M-0072026-08-01instance_appendedIN-006IN-007
M-0062026-07-09description_reorderedDESCRIPTION-REORDERED
M-0052024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
7 instances on recordShow sources ↓Hide ↑
IN-001GitHub Copilot productivity study — Peng et al. (Microsoft Research)1. Peng, S., Kalliamvakou, E., Cihon, P., and Demirer, M. (2023), The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, arXiv:2302.06590. · Study Design; Results; Discussionsupportive
IN-002Noy and Zhang — Experimental evidence on productivity effects of generative AI in professional writing1. Noy, S. and Zhang, W. (2023), Experimental evidence on the productivity effects of generative artificial intelligence, Science 381(6654), 187–192. DOI 10.1126/science.adh2586 · Abstract and experimental resultssupportive
IN-003AI in legal practice — contract review and due diligence deploymentsupportive
IN-004Hallucination and reliability failure documentation — legal and medical contexts1. Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023), Opinion and Order on Sanctions. · Opinion and Order on Sanctions, 22 June 20232. Omiye, J. A. et al. (2023), Large language models propagate race-based medicine, npj Digital Medicine 6, 195. DOI 10.1038/s41746-023-00939-z · Abstract; Results; Discussioncontesting
IN-005Brynjolfsson, Li, and Raymond — Generative AI at work (customer service study)supportive
IN-006Agentic deployment failures — early autonomous task completion attemptspartial
IN-007ORCA-bench — live-system root-cause analysis under limited supervisionCONTESTING