← ObservatoryThe RecordFR-AI-0009
PROG-AI
FR-AI-0009

World Models — Physical Prediction and Transfer

AI systems can learn predictive representations of the physical world that support reliable action when the objects, environment, task, or embodiment differ materially from those encountered during training.

EscalatingVS-02·since 2026-08-19
Verification Matrix

Verification position derived from the record’s assessments; dates show when Faultline first recorded each stage.

VS-01
Assertion
VS-02
Published evidence
Current from 2026-08-19 — present
VS-03
Audit
VS-04
Replication
VS-05
Operation
Stage first recorded Current verification position Not yet recorded
State Warrant
Current stateEscalatingVS-02
Why this state?Admission evidence basis: Hafner et al., Nature 640 (2025), 'Mastering diverse control tasks through world models'; Assran et al., arXiv:2506.09985 (2025), 'V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning'; Google DeepMind, 'Genie 3: A new frontier for world models' (2025); Yang et al., arXiv:2605.29360 (2026), 'MiraBench'; Cai et al., arXiv:2605.27589 (2026), 'What-If World'; Jiang et al., arXiv:2604.19092 (2026), 'RoboWM-Bench'; Joseph et al., ICML / arXiv:2602.07050 (2026), 'Interpreting Physics in Video World Models'. Formal admission and material commitments recorded in Drive on 2026-08-19.
Assessment summaryThe claim enters the corpus under genuine two-sided pressure. V-JEPA 2-AC (IN-002) is substantive supportive evidence: after large-scale observational pretraining and limited robot-video adaptation, a learned action-conditioned model supported zero-shot planning on Franka arms in two target laboratories without target-environment robot data or task-specific reward. DreamerV3 (IN-001) independently shows broad world-model control generality across more than 150 tasks, and Genie 3 (IN-003) shows that interactive action-responsive simulation has advanced beyond passive video generation. Those results do not settle the class-level claim. MiraBench (IN-004), What-If World (IN-005) and RoboWM-Bench (IN-006) directly expose the central boundary: visually plausible futures can be wrong about commanded actions, causal interventions, contact dynamics and executable behaviour. The Pressure State is ESCALATING because credible positive capability and credible failure evidence are both strengthening. Verification Stage is VS-02 because bounded demonstrations and dedicated challenge benchmarks now exist, but reliable transfer across materially different tasks and embodiments has not yet received sufficiently broad independent audit or operational replication.
In this state since2026-08-19
Mechanisms

Causal mechanisms recorded for this claim. The State Warrant above remains the authoritative current assessment.

BOTTLENECK — MEASUREMENT VALIDITYBN-001

Visual plausibility as a proxy for action-conditioned and physical fidelity. Conventional video-generation metrics and human realism judgments can rate a future as convincing even when it responds incorrectly to the action or intervention that produced it. MiraBench (IN-004) and What-If World (IN-005) show that this proxy gap is operationally important. The record cannot advance on photorealism alone; evidence must increasingly test whether predicted consequences are causally and physically appropriate for the action taken.

Resistance MechanismRM-001

Distribution and embodiment shift. Predictive representations can encode camera geometry, local object distributions, robot morphology and interaction statistics that are stable inside the training distribution but change under deployment. IN-002 demonstrates some environment transfer, but the strongest positive evidence remains within a constrained manipulation setting and one general robot family. Cross-task and cross-embodiment reliability therefore remain the principal generalisation boundary.

Resistance MechanismRM-002

Compounding prediction error under extended interaction. Autoregressive and recurrent world models repeatedly condition later predictions on earlier predicted states, allowing small spatial, contact or causal errors to accumulate. Genie 3 (IN-003) explicitly identifies interaction duration as bounded, while RoboWM-Bench (IN-006) documents local physical inconsistencies that can become execution failures. Long-horizon reliability is therefore harder than short-horizon visual coherence.

AttractorAT-001

Scalable observational pretraining plus sparse action-conditioned adaptation. V-JEPA 2-AC (IN-002) suggests that very large observational video corpora can supply reusable physical priors which comparatively small quantities of robot interaction data can convert into planning capability. If this pattern replicates across substantially different tasks, environments and embodiments, it would provide a scalable route around the cost of collecting task-specific physical interaction data for every deployment setting.

Assessment History
2026-08-19
Initial assessment — Escalating
The claim enters the corpus under genuine two-sided pressure. V-JEPA 2-AC (IN-002) is substantive supportive evidence: after large-scale observational pretraining and limited robot-video adaptation, a learned action-conditioned model supported zero-shot planning on Franka arms in two target laboratories without target-environment robot data or task-specific reward. DreamerV3 (IN-001) independently shows broad world-model control generality across more than 150 tasks, and Genie 3 (IN-003) shows that interactive action-responsive simulation has advanced beyond passive video generation. Those results do not settle the class-level claim. MiraBench (IN-004), What-If World (IN-005) and RoboWM-Bench (IN-006) directly expose the central boundary: visually plausible futures can be wrong about commanded actions, causal interventions, contact dynamics and executable behaviour. The Pressure State is ESCALATING because credible positive capability and credible failure evidence are both strengthening. Verification Stage is VS-02 because bounded demonstrations and dedicated challenge benchmarks now exist, but reliable transfer across materially different tasks and embodiments has not yet received sufficiently broad independent audit or operational replication.
Admission evidence basis: Hafner et al., Nature 640 (2025), 'Mastering diverse control tasks through world models'; Assran et al., arXiv:2506.09985 (2025), 'V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning'; Google DeepMind, 'Genie 3: A new frontier for world models' (2025); Yang et al., arXiv:2605.29360 (2026), 'MiraBench'; Cai et al., arXiv:2605.27589 (2026), 'What-If World'; Jiang et al., arXiv:2604.19092 (2026), 'RoboWM-Bench'; Joseph et al., ICML / arXiv:2602.07050 (2026), 'Interpreting Physics in Video World Models'. Formal admission and material commitments recorded in Drive on 2026-08-19.
Claim Lineage

Historical narrative recorded for this claim. It does not override the current State Warrant.

2018–22
Learned world models become an explicit control paradigm. Latent-dynamics approaches show that agents can learn compact predictive states and plan or learn policies inside them, establishing the architectural idea without settling broad transfer.
2023–25
DreamerV3 demonstrates unusually broad task coverage under one world-model reinforcement-learning configuration. World-model capability expands across control benchmarks, but most evidence remains inside designed environments.
2025
The frontier moves toward observational pretraining and physical planning. V-JEPA 2-AC demonstrates bounded zero-shot planning in new laboratory environments; Genie 3 demonstrates real-time interactive generated worlds. The claim becomes empirically testable beyond passive prediction.
2026
Dedicated causal and embodiment-grounded benchmarks expose a reliability gap. MiraBench, What-If World and RoboWM-Bench show that visual realism does not reliably imply action fidelity, causal correctness or physically executable behaviour. The frontier shifts from generating plausible worlds to predicting consequences that remain useful under intervention and transfer.
Open Questions

Questions retained in this record. The current State Warrant may have narrowed or reframed earlier questions.

OQ-001

What minimum change in object distribution, environment, task or embodiment should count as material transfer rather than interpolation within the training distribution?

Raised 2026-08-19
OQ-002

Can action-conditioned world models become physically reliable without explicit factorised physics variables, or do distributed learned representations create failure modes that only become visible under intervention?

Raised 2026-08-19
OQ-003

Once dedicated causal and execution benchmarks are used, does improvement in generated-world visual fidelity correlate meaningfully with action-conditioned and physically executable fidelity?

Raised 2026-08-19
OQ-004

Can a world model trained largely from observation transfer useful planning capability across materially different robot embodiments without extensive new interaction data?

Raised 2026-08-19
OQ-005

What evidentiary threshold should be required before success in synthetic or generated environments is treated as evidence of reliable action in the corresponding real physical environment?

Raised 2026-08-19
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0082026-09-07provenance_reviewLPR-001-D09
M-0072026-08-28reference_correctedInstance-level references absentIN-001–IN-008 source references recorded
M-0062026-08-22instance_addedIN-001–IN-007IN-001–IN-008
M-0052026-08-19diagnosis_heldDIAGNOSIS-HELD
M-0042026-08-19mechanisms_recordedMECHANISMS-RECORDED
M-0032026-08-19assessment_issuedASSESSMENT-ISSUED
M-0022026-08-19instances_loggedINSTANCES-LOGGED
M-0012026-08-19record_createdRECORD-CREATED
Evidence Sources
8 instances on recordShow sources ↓Hide ↑
IN-001DreamerV3 — general world-model control across more than 150 tasks1. Hafner, D. et al. (2025), Mastering diverse control tasks through world models, Nature 640, 647–653. · Abstract; evaluation across eight domains and more than 150 tasks; fixed-hyperparameter comparisonpartial
IN-002V-JEPA 2-AC — zero-shot robot planning in two new laboratory environments1. Assran, M. et al. (2025), V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning, arXiv:2506.09985. · Abstract; >1 million hours pretraining; <62 hours DROID post-training; zero-shot Franka deployment in two labssupportive
IN-003Genie 3 — real-time interactive generated worlds with bounded consistency1. Google DeepMind (2025), Genie 3: A new frontier for world models. · Capabilities; environmental consistency; limitations; 720p, 24 fps, several-minute interactionpartial
IN-004MiraBench — visual fidelity fails as a proxy for action-conditioned reliability1. Yang, T. et al. (2026), MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models, arXiv:2605.29360. · Abstract; >16,000 judgments; 12 model configurations; three central findingscontesting
IN-005What-If World — causal intervention benchmark exposes systematic failures1. Cai, K. et al. (2026), What-If World: A Causal Benchmark for General World Models in Embodied Scenarios, arXiv:2605.27589. · Abstract; 319 intervention pairs; nuScenes and DROID; nine-model paired-score resultscontesting
IN-006RoboWM-Bench — generated manipulation behaviours remain difficult to execute physically1. Jiang, F. et al. (2026), RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation, arXiv:2604.19092. · Abstract; embodied-action conversion and execution; failure modes; fine-tuning resultcontesting
IN-007Mechanistic evidence — physical variables emerge in distributed video-model representations1. Joseph, S. et al. (2026), Interpreting Physics in Video World Models, arXiv:2602.07050. · Abstract; Physics Emergence Zone; speed, acceleration and motion-direction probes; distributed representation conclusionpartial
IN-008CaliBench — stochastic physical outcomes remain systematically miscalibrated1. Sadeghi, J. et al. (2026), CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?, arXiv:2608.16829. · Abstract; nine scenes; six image-to-video models; calibration and probability-mass concentration resultscontesting