Hao Liu
Co-founder of Shiqi Future, Head of World Model & Embodied Intelligence
Hao Liu is the Co-founder of Shiqi Future and Head of World Model & Embodied Intelligence. He developed the industry's first universal grasping foundation model for open-world objects (DINO-X-Grasp) and led the creation of an industry-leading unified world model that integrates both generation and understanding. He previously served as the Head of Autonomous Driving Algorithm for Mass Production at Didi, where he delivered the industry's first mass-produced integrated cockpit–driving algorithm (Xiaopeng MONA). Prior to that, he was the Head of L4 Perception at WeRide, responsible for the industry's first low-compute LiDAR-based multimodal perception algorithm, as well as the AI inference framework and middleware for WeRide from scratch. Earlier in his career, he served as the Head of AI Infrastructure at Baidu, where he developed the AI inference framework Anakin, the distributed training communication library for PaddlePaddle, and Baidu's unified AI heterogeneous computing cluster. In 2017, he was a recipient of Baidu's highest award—the Million Dollar Award—as a member of the Kongming Team.
Topic
World Models Need to Understand Objects and Actions: Unified Architectures and Embodied AI in Practice
The goal of a world model is to learn the laws of physics and the principles of causal interaction. Seequ believes that all physical laws act upon objects; therefore, a world model must first understand objects before it can learn physical laws. Similarly, causal interactions in embodied AI are fundamentally driven by actions and their consequences, meaning that a world model must first understand actions before it can learn causal interaction patterns. Based on Seequ’s unified “Generation + Understanding” world model—DINO-X, a vision-native foundation model—this session will share the latest progress in learning physical laws and causal interaction patterns. Drawing on the real-world deployment of a general-purpose robotic grasping system, the talk will also explore how first-principles thinking can help address “physical hallucinations” in AI models, along with the architectural design and engineering optimization practices behind it. Outline 1. Current Challenges: Why Should World Models Be Vision-Native? A brief review of the evolution of world models and an exploration of why video generation alone is insufficient for truly learning physical laws. 2. First Principles: Why Must World Models Understand Objects and Actions First? An examination of why object understanding is fundamental to learning physical laws, and why action understanding is essential for modeling causal interactions in embodied AI. 3. Architecture Deep Dive: The Unified “Generation + Understanding” World Model An overview of the core network architecture and training data construction practices behind Seequ’s DINO-X. 4. Embodied AI in Practice and Engineering Optimization An exploration of multi-view capabilities in complex physical interactions, as well as inference framework deployment and optimization practices in conjunction with underlying hardware. 5. Case Studies and Lessons Learned Live demonstrations of video generation and embodied robotic grasping, along with typical bad cases and practical lessons for avoiding common pitfalls. Key Takeaways Gain a first-principles understanding of how world models can learn physical laws and causal interaction patterns, and develop a clearer view of the frontier and future trends of world models. Learn the design principles behind unified multimodal architectures that combine generation and understanding. Gain practical experience in deploying cutting-edge models for embodied AI applications, including generalizable robotic grasping, from training through inference and real-world deployment.