World Model Advances: Multiple Institutions Release New Results, Accelerating from Generation to Action
During and around WAIC 2026, multiple institutions intensively released world model-related achievements, covering foundation models, embodied intelligence, digital content, and industrial applications, marking the accelerated transition of world models from academic concepts to industrial deployment.
Kunlun: Full-Modal Model Matrix, Proclaiming "Year of the World Model"
Kunlun unveiled a full-modal model matrix spanning the physical world, foundation layer, and digital content layer at WAIC. CEO Fang Han declared 2026 as the "Year of the World Model."
- Riemann-1.0: An embodied action model adopting the "world action model" approach, unifying robot action generation and environment prediction into a single framework. It achieves a 99.0% success rate on the LIBERO benchmark, 62.6% on RoboCasa-365, and 94.3% on RoboTwin2.0. After pretraining with first-person human data, long-horizon seen-scene success rates improve by 28.15%, and unseen scenes by 16.97%.
- Matrix-Game 3.5: A real-time interactive world model supporting continuous interaction and long-term spatial memory, capable of generating structurally self-consistent open worlds.
- Mureka V9.5 and O3: Applying cognitive and reasoning capabilities to music creation.
Jijia Vision: General World Model Product Matrix, from Content Creation to Embodied Intelligence
Jijia Vision showcased a four-in-one system at WAIC: "World Model - Embodied Foundation Model - Native Embodiment - Generalization Scenarios":
- World Generation Models: YiSu (content creation), DriveDreamer (autonomous driving simulation), GigaWorld (embodied intelligence training and validation).
- World Action Models: GigaBrain (general embodied brain), GigaWorld-Policy (direct action output).
- Embodiment Products: Shiguang S1 (home scenarios), Maker H01 (smart manufacturing), demonstrating long-horizon real-world operations.
Yann LeCun's Team: AdaJEPA — Test-Time Adaptive World Model
Yann LeCun's team proposed AdaJEPA, achieving the first test-time adaptation of a JEPA world model within an MPC planning loop. Core innovations:
- Self-supervised fine-tuning using state transitions from the agent's own interactions, requiring no external labels.
- Only a single gradient update per re-planning, fine-tuning only a few top layers of the encoder and predictor, adding 0.01–0.03 seconds of latency per step.
- On PushT and PointMaze benchmarks, out-of-distribution task planning success rates improve significantly, with the largest gains when training data is scarce.
LingBot World 2.0: Open-Source Infinite Interactive World Model
LingBot-World-Infinity (14B main model + 1.3B distilled model) achieves four innovations:
- MoBA Hybrid Attention + Flow Matching Pretraining: Continuously generates stable output for 1 hour without quality degradation.
- Consistency Distillation + Distribution Matching Distillation: Supports real-time interaction at 720P/60fps.
- Multi-Granularity Hierarchical Data Engine: Fuses real footage, game synthesis, and web videos with segmented temporal annotations.
- Director-Navigator Dual-Agent Scheduling: Supports diverse interactions (melee, spellcasting, weather switching) and multiplayer online play.
DAMO Academy: RynnWorld-Teleop and RynnWorld-4D
Alibaba DAMO Academy released two works:
- RynnWorld-Teleop: A digital teleoperation solution replacing real robots with generative world models. Operator gestures drive real-time video generation, automatically obtaining joint-level action labels. Real-robot experiments achieve zero-shot Sim2Real transfer; synthetic data combined with real data stably improves success rates. Code and models are open-sourced.
- RynnWorld-4D: A 4D embodied world model using RGB+Depth+Optical Flow (RGB-DF) representation, simultaneously generating future RGB video, depth maps, and optical flow. It constructs the Rynn4DDataset 1.0 with 254.4 million frames. The accompanying policy head, RynnWorld-4D-Policy, achieves 9Hz equivalent closed-loop control, achieving SOTA on bimanual dexterous manipulation tasks.
Zhongshu Ruizhi: Causal World Model Deployed in 35 State-Owned Enterprises
Zhongshu Ruizhi released the "AI for Reasoning" causal intelligence system and the "Causal World Model Technical System Blue Book." Core features:
- Based on the "meta-causality" concept, dynamically constructing causal graphs supporting intervention and counterfactual reasoning.
- In oil and gas drilling well control scenarios, it advances warning windows by approximately 15–20 minutes, achieves root cause localization accuracy of ~94%, and filters over 40% of invalid false alarms.
- Deployed in over 35 large state-owned enterprises, covering more than 800 scenarios, with cumulative safe operation exceeding 15,000 hours.
Also available in 中文.