World Models Become the New Frontier in Embodied AI

Avatar 0

Staff Reporter | Xu Meihui

News Editor | Wen Shuqi

Heading into 2026, the robotics industry’s focus is rapidly shifting from the physical body to the “brain,” and world models are suddenly the hottest topic in town.

Lyu Yao, Chief Scientist at Guangxiang Lab, recently told reporters at NUPIAO that judging a world model isn’t just about whether the next frame looks realistic. The real test is whether the model can correctly explain why the world changes when an action changes. In his view, only by learning the underlying causal laws of the physical world can we truly achieve general-purpose embodied intelligence.

But even after the “brain” learns to understand the world, there’s still no industry-wide consensus on what architecture robots should use to actually learn how to act.

In the embodied AI space, the discussion around robot “brains” largely revolves around two technical approaches: VLA and world models. They aren’t strictly mutually exclusive, but they have different priorities:

One path is VLA (Vision-Language-Action models), which connects visual and language inputs directly to robot actions. Typical models learn from robot demonstration data to figure out how to generate actions based on what they observe and what they’re told to do.

The other path is the world model approach. Whether it’s generative models like Sora, spatial intelligence models like those from World Labs, or visual representation models like JEPA, they all follow the same fundamental modeling paradigm: learning statistical representations of the world from massive amounts of data, then predicting future changes from observations.

The debate over which route to take isn’t new this year. Back at the 2025 World Robot Conference, Wang Xingxing, founder of Unitree Robotics, publicly questioned the then-popular VLA approach.

He argued that the biggest bottleneck for mass adoption of humanoid robots lies in AI models and, more specifically, in model architecture, not in hardware or data as the industry tends to focus on. He even called VLA a “relatively foolproof but simplistic architecture.” In contrast, he favored driving robots through video generation models or world models, saying this path has a “higher probability of convergence.”

A year later, world models have clearly gained momentum, but the technical debate hasn’t settled down.

On the VLA front, Lyu Yao believes it’s hit a ceiling. The front end is a language model, and the back end is an action expert head—there’s a fundamental modality gap when mapping language tokens to action spaces. He thinks it’s basically unrealistic to train VLA models on super-large-scale data to achieve truly generalizable policies.

Lyu says that whether companies are pivoting to world models because VLA competition got too fierce, or because they hit VLA’s limits during internal testing, the industry is already rethinking its original approach.

But that’s about where the consensus ends. The concept of “world model” still lacks a unified definition to this day.

The most obvious sign of this definitional chaos? People are talking about completely different things when they say “world model.” Fei-Fei Li and her team at World Labs wrote a paper this year specifically tackling this concept, categorizing world models by function into three types: renderers, simulators, and planners. They noted that “world model” has become one of the most important—and most misused—terms in AI today.

Under this framework, renderers mainly generate pixel-level visuals for humans to watch, simulators require more accurate descriptions of geometry, physics, and dynamic changes, and planners are responsible for deciding the next action based on observations and goals. Interestingly, World Labs also classifies VLA under the “planner” category, which suggests VLA and world models aren’t necessarily an either-or proposition.

The disagreement isn’t just about definitions either. Guo Yandong, founder and CEO of Alpha Robotics, said at the Beijing BAAI Conference this June that world models aren’t a competing route to VLA—they’re actually a core component within the VLA framework.

Beyond the architectural debates, there are some very practical problems. One company executive working on world models told reporters at NUPIAO that some video generation results look impressive at first glance, but they still fall short when it comes to accurately expressing physical properties like friction. When you actually try to use them for robot manipulation tasks, they’re just not good enough. Some of the necessary data still has to be collected by companies themselves, which doesn’t come cheap.

The same executive told us that the industry hasn’t even agreed on a standard for how to grade robot capabilities. Right now, robots can handle some well-defined, short-process tasks in factories, but if you switch to something open-ended like cleaning a house, they need to not only recognize the environment but also understand goals and causal relationships, and continuously adjust their actions based on results. We’re still a long way from truly pulling that off.

Guangxiang is taking a different angle by tackling the question: “How will actions change the world?” Recently, Guangxiang Technology, in collaboration with Professor Li Shengbo’s research group at Tsinghua University, released its first-generation physics-native world model, Phi-WM 1.0 ActEffect. The core idea is to turn the world model into a “training-time feedback mechanism” that teaches robots what consequences their actions will bring.

Specifically, ActEffect can generate three action plans at different levels of detail in one go, predicting the likely outcome of each plan, and then comparing them to pick the best one. Once training is complete, the world model is removed from the inference chain, and at deployment, the policy network directly outputs actions in a single forward pass.

In Lyu’s explanation, it’s like “distilling” the world model’s judgment about action outcomes directly into the policy weights ahead of time, which reduces the computational burden during deployment.

Image Source: VCG

Beyond models, data is another unavoidable issue. Whether world models can actually reduce the data requirements for robots is another question the industry can’t agree on.

Regarding the claim that “world models will reduce data needs,” Lyu thinks it needs to be discussed case by case. He doesn’t buy the idea that pure video-prediction world models can naturally cut down data requirements. Training a video prediction model that’s sufficiently realistic and generalizable still requires massive amounts of data itself, and then using it to generate data to fill gaps isn’t necessarily a sound approach.

“But if a world model can distill cross-scenario, universal physical laws from data, then it could potentially reduce the need for cross-scenario and cross-platform data,” Lyu said.

Zhang Tao, founder and CEO of Guangxiang Technology, offered a relatable example: if you throw ten objects of different weights and shapes into a basket, and you’ve practiced ten times and can make all ten in, you still have to start from scratch when you’re handed an eleventh or hundredth object you’ve never seen. Following this logic, what’s most scarce isn’t successful demonstrations—it’s data from massive trial-and-error and counterfactual reasoning in simulation.

On this point, Lyu estimates that the counterfactual trial-and-error data needed could be 10 to 100 times more than real successful demonstrations. But when it comes to real-robot post-training at a specific workstation, the data volume is roughly a dozen to a few dozen hours, at most not exceeding 100 hours. Following Guangxiang’s approach, the world model pushes more of the trial-and-error into lower-cost simulation environments, reducing the need for some of the expensive real-robot data.

But factories aren’t just looking at how much data a model needs. When it comes to commercial deployment, the battle over technical routes ultimately becomes a battle over ROI.

Research released by IDC this August shows that China’s industrial embodied AI robot market is expected to be around 5.74 billion yuan in 2025. Their user survey found that over 80% of manufacturing companies want their industrial embodied AI robot projects to achieve payback within two years.

What this means is that for industrial clients, it doesn’t really matter whether it’s VLA or world models. What matters is whether it can be deployed quickly, run stably, and pay for itself within the expected timeframe.

Even though the technical route hasn’t converged, that hasn’t stopped the money from pouring in. CB Insights data shows that global investment in world models jumped from $1.4 billion in 2024 to $6.9 billion in 2025—nearly a fivefold increase.

And the big funding rounds have continued into 2026. In February, Fei-Fei Li’s World Labs announced a new $1 billion funding round. In March, AMI Labs, co-founded by Yann LeCun, closed a $1.03 billion round. World Labs is focused on spatial intelligence and world models, while AMI Labs is centered on modeling, reasoning, and planning for the real world.

Capital has already placed its bets, but robots aren’t making their way into real-world scenarios quite as quickly.

“The embodied AI industry is very likely to go through a major shakeout, just like autonomous driving did back in the day,” Lyu told reporters at NUPIAO, predicting that the capital markets will eventually return to rationality, and companies lacking a core technical moat and genuine delivery capabilities will be the first to be weeded out.

That shakeout might not be far off. Zhang Tao believes that if the industry stays in its current state, a cold winter could indeed hit within one to two years—but the outcome isn’t set in stone. In his view, the real wildcard is whether we can see the first market-validated embodied AI application with the potential to scale and create real commercial value before that happens.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter code for secure login, or use password.

Code Login Password Login