中文
← Back to news
ModelsJul 20, 2026

WAIC 2026: World Models and Physical AI Advance with Multiple Technical Routes and Accelerated Industrial Deployment

At the 2026 World Artificial Intelligence Conference (WAIC), world models emerged as one of the most prominent directions. Multiple companies showcased complete technology stacks ranging from digital content generation to physical world interaction, signaling the transition of this field from academic concepts to industrial deployment.

Core Trend: From Generation to Action, World Models Move Toward Integration

Current world model development features a landscape of parallel technical routes. Representational, generative, and interactive models respectively address parts of world understanding, state inference, and environmental interaction, but the industry is accelerating toward an integrated understanding-generation-prediction action world model. Daxiaorobot released Kairos 3.1, which integrates generative intelligence, physical intelligence, and cognitive intelligence, employing a hybrid Transformer architecture to compress multi-source information into a unified latent space, achieving native unification of understanding, generation, and prediction. Kunlun Tech's Riemann Dynamics introduced Riemann-1.0, representing the third-generation route WAM (World Action Model), which incorporates action generation and environmental state evolution into a single causal generation framework.

Key Breakthroughs: Long-Term Temporal Stability and Real-Time Interaction

Interactive world models have long faced two major bottlenecks: long-term temporal error accumulation and high computational cost for high-fidelity interaction. LingBot-World-Infinity 2.0 achieves stable generation for up to one continuous hour without quality degradation through its self-developed MoBA mixed attention and flow matching pre-training, and distills a 1.3B lightweight model supporting 720P/60fps real-time interaction. Alibaba DAMO Academy's RynnWorld-Teleop uses streaming autoregressive distillation to boost video generation speed from 2.8FPS to 40FPS, meeting real-time teleoperation requirements.

Data-Driven: Human Videos Become Key for Training

Scarcity of high-quality data is a core bottleneck for embodied intelligence. Multiple companies have validated the effectiveness of the human first-person video pre-training + small robot data fine-tuning paradigm. Riemann-1.0 uses 232,000 hours of training data (200,000 hours of which are human videos), achieving a 62.6% success rate on the RoboCasa-365 benchmark, an 8.4 percentage point improvement over the previous SOTA. Daxiaorobot proposed the Information Density Law, categorizing embodied data into five levels L1-L5, and released an environment-based data collection solution 2.0, with its ACE Sense Glove achieving force-tactile sensitivity of 0.01N. RoboScience's Visics model reduces per-data cost to 1/20 to 1/200 of traditional solutions through an automated annotation pipeline.

Industrial Deployment: From Lab to Real-World Scenarios

World models are accelerating into practical applications. Daxiaorobot's fulfillment robot W1 is already operating in retail scenarios such as Shaomai Gou and Kuaikeda, with plans to deploy in 1,000 locations over the next year. Jijia Shijie showcased a complete product matrix covering home service (Shiguang S1) and intelligent manufacturing (Maker H01). Zhongshu Ruizhi's causal world model has been deployed in over 35 large state-owned enterprises, covering more than 800 scenarios including oil and gas drilling and power dispatch, with cumulative safe operation exceeding 15,000 hours.

Delivery Model Innovation: Cloud-Based and Cross-Embodiment Generalization

The delivery model for embodied intelligence is shifting. RoboScience, in collaboration with Tencent Cloud, launched the world's first cloud-based embodied large model Visics, supporting EaaS (Embodied AI as a Service) mode, enabling 30-second handover and zero-shot generalization across over 10 dexterous hand types. Kunlun Tech's Riemann-1.0 also emphasizes cross-embodiment capabilities, covering 41 robot embodiments. This approach of decoupling models from hardware is expected to reduce the cost of large-scale replication of embodied intelligence.

Also available in 中文.