Reinforcement learning from human feedback produces AI systems whose behaviour continues to reflect human preferences when deployed beyond the conditions represented in training.
Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.
Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.
Reward model distributional brittleness. RLHF trains a reward model on human preference data and then optimises the language model against that reward model. The reward model is itself a learned approximation of human preferences, trained on a finite distribution of examples. When the language model encounters inputs outside that distribution, the reward model's approximation degrades — it assigns high reward to outputs that humans in novel contexts would not prefer. This is the structural source of jailbreak and sycophancy failures: the reward model cannot accurately represent preferences it was not trained to evaluate. The mechanism is not a failure of RLHF as a concept; it is a limitation of any finite training distribution.
Preference proxy misalignment. Human raters during RLHF training evaluate outputs on dimensions they can perceive and assess — fluency, apparent helpfulness, surface agreement with their views. They cannot reliably rate outputs on dimensions that require expertise they lack, careful verification they do not have time for, or long-horizon consequences they cannot observe. The training signal therefore reflects what raters can easily evaluate rather than what they would prefer if fully informed. This produces systematic gaps between trained preferences and actual preferences, which widen as the model operates in novel contexts where the gap between easy-to-rate and actually-preferred is larger than in training.
No agreed measurement of preference reflection under distribution shift. The claim requires that preference reflection be measurable outside training conditions. No standard evaluation exists for this. Red-teaming measures adversarial robustness but not general generalisation. Human evaluation measures perceived quality in evaluated contexts but not behaviour in unevaluated contexts. Interpretability methods can identify some internal representations but cannot directly measure preference generalisation. Without a measurement instrument, the claim cannot transition from FRAGMENTING to any resolved state regardless of training improvements. The bottleneck is metrological: the thing the claim asserts cannot currently be measured directly.
Interpretability-grounded preference verification. The scalable oversight and interpretability research programmes (INST-006) represent a potential resolution path: if internal model representations of human preferences can be identified and verified, preference generalisation can be assessed directly rather than through behavioural proxies. This is the attractor for this record — not better RLHF training, but a measurement capability that would allow the claim to be evaluated rather than merely approximated. The attractor is early-stage; current interpretability tools cannot yet verify preference representations reliably. But its direction is visible, and it is the research trajectory most likely to resolve BN-001.
Historical narrative recorded for this claim. It does not override the current State Warrant.
Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.
Does increasing model capability make preference generalisation better or worse? INST-005 suggests worse; INST-006 suggests potentially better under specific training regimes. This question is not resolvable from current evidence and is the programme's central unresolved tension.
Raised 2024-01-15Can interpretability tools eventually provide direct measurement of preference representation in model weights, resolving BN-001? If so, the claim becomes evaluable rather than merely approximable. If not, the claim may be permanently unresolvable as stated.
Raised 2024-01-15The three failure modes (adversarial, sycophantic, capability-outpacing) are independent. Does solving one have any effect on the others, or do they require independent solutions? Current evidence does not address this.
Raised 2024-01-15Is this a class-level claim that should remain at the level of RLHF as a method, or should separate records track specific training regimes (standard RLHF, Constitutional AI, RLAIF) as the methods diverge? The corpus lesson from FR-QE-0002 applies: bundling claims that resolve on different timescales produces bottlenecks that belong to the claim rather than the frontier.
Raised 2024-01-15| Mutation | Date | Field | Prior value | Current value |
|---|---|---|---|---|
| M-013 | 2026-09-06 | description_restored | Legacy ingestion cutoffs: mechanisms:RM-001, mechanisms:RM-002, mechanisms:BN-001, mechanisms:AT-001, lineage:2017–20 | Source-restored complete descriptions |
| M-012 | 2026-09-01 | assessment_correction | AS-001 / AS-002 source-dependent wording | AS-003 |
| M-011 | 2026-09-01 | provenance_correction | LPR-001-D03 discrepancies_found | LEGACY-INSTANCES-CORRECTED-ASSESSMENT-REVIEW-PENDING |
| M-010 | 2026-09-01 | provenance_review | — | LPR-001-D03 |
| M-009 | 2026-07-14 | assessment_issued | AS-001 | AS-002 |
| M-008 | 2026-07-14 | instances_appended | — | IN-007 / IN-008 / IN-009 |
| M-007 | 2026-07-09 | description_reordered | — | DESCRIPTION-REORDERED |
| M-006 | 2024-01-15 | programme_panel_added | — | PROGRAMME-PANEL-ADDED |
| M-005 | 2024-01-15 | null_condition_result | — | NULL-CONDITION-RESULT |
| M-004 | 2024-01-15 | mechanisms_recorded | — | MECHANISMS-RECORDED |
| M-003 | 2024-01-15 | assessment_issued | — | ASSESSMENT-ISSUED |
| M-002 | 2024-01-15 | instances_logged | — | INSTANCES-LOGGED |
| M-001 | 2024-01-15 | record_created | — | RECORD-CREATED |