Reporter |
Editor | Liu Fangyuan
Ant Lingbo first caught the public’s eye back in February 2025. A flurry of job postings suddenly put this embodied intelligence company under the spotlight. Since then, they’ve been churning out research results left and right, and their game plan is finally coming into focus.
Just last Friday, Ant Lingbo dropped the industry’s first-ever embodied-native world action model, LingBot-VA 2.0. Thanks to its native architecture, this model shows some seriously impressive speed and adaptability in real-world tests. For instance, the robot can go toe-to-toe with a human in multiple rounds of impromptu table tennis, all without relying on any external cameras.

LingBot-VA 2.0 isn’t just a product; it’s a statement. It represents a key strategic choice: start from scratch and pre-train on an autoregressive architecture. At its launch, Ant Lingbo’s CEO, Zhu Xing, and Chief Scientist, Shen Yujun, sat down with the press for the first time to spill the beans on their tech roadmap and big-picture strategy.
Ant Lingbo’s mission in the embodied intelligence ecosystem is clear: be the universal brain company. Their job is to crack the generalization problem for robots. “We decided way back in early 2025 that we’d go the ‘brain’ route,” Shen Yujun told us in an interview. “We figured that’s the biggest missing piece for getting robots out of the lab and into the real world.”
Shanghai Ant Lingbo Technology was officially founded at the end of 2024, but the team, led by Shen, had already been tinkering with embodied intelligence tech for a while before that.Zhu Xing summed it up: “In the first half of 2025, our focus was on building the team and getting our data ready.”Back in January, they released their first open-source spatial perception model, LingBot-Depth 1.0. Then, the big push was getting the model into production. With the launch of 2.0, they’ve finally coined the term “embodied-native.”

So, what’s this “embodied-native” road all about? It means they’re fully committed to training models from scratch, completely on their own terms, without leaning on capabilities built for the digital world.
This approach sets them apart from most other players in the field. “The digital world has a lot of demands that just don’t apply in the physical world,” they argue. “So, we’re taking what works from digital models, boosting the stuff that matters for physical tasks, ditching what doesn’t, and then totally re-architecting the model. We’re training it from the ground up using a mix of internet data and real robot data.”
The embodied intelligence industry is still in its infancy, and the tech paths for building the brain are far from settled. The two main schools of thought are VLA (Vision-Language-Action models) and VA (Vision-Action models), or even WAM (World-Action models). The first one gets a lot of love because it’s built on multimodal models that can understand human intent, and it’s less resource-hungry for reasoning, which means lower costs to deploy.
But starting around March last year, the Ant Lingbo team realized that multimodal models have a blind spot: they’re not great at prediction. And since robots have to get things done in the physical world, prediction is a big deal. So, they set out to fix that, and by January this year, they’d released their VA (Vision-Action) model. By adding dynamic modeling, the model could start to predict what’s coming next.
Looking at the current landscape, Ant Lingbo’s core beliefs are two-fold. First, models built for the digital world aren’t naturally a great fit for robots. Second, they want to design a model that’s tailor-made for robots and train it up. The goal is to give the robot both the ability to understand and the ability to generate, but right now, there’s no open-source model in the digital world that can serve that purpose. “That’s the biggest difference between us and everyone else,” Shen Yujun stated.
Because of this, Ant Lingbo doesn’t have a strong preference for one tech route over the other. In fact, they believe that in the future, these two paths are bound to converge, creating a “1+1>2” model that mixes VA and VLA. How exactly to make that happen? They’re still figuring that out.

Once the model architecture logic is set, the foundational problem Ant Lingbo needs to tackle is data. This is widely recognized as the bottleneck holding back embodied intelligence. Zhu Xing emphasized to the press, “If your data doesn’t take off—in terms of scale, quality, or distribution—then your model architecture is just a castle in the sky.”
But, of course, this is a long-term game that demands serious investment.
The embodied-native route means Ant Lingbo needs both product data and custom-collected data. Their philosophy? They trust the value of real-world data more than anything.
Zhu Xing noted that data collection methods are evolving fast. “Looking at the shifts in collection schemes like Ego and UMI recently, the efficiency is skyrocketing, which means the cost of real robot data is dropping quickly. But from the perspective of what physical intelligence needs, we’re still way short of the mark.” He added that the model’s own application and iteration in the real world is critical. “To truly make the data flow, we need real feedback from real-world applications.”
That’s what’s behind Ant Lingbo’s flurry of activity this year. We’ve learned that the company is teaming up with partners like Jianzhi Technology in a data alliance to build a standardized data system.
Being under the Ant Group umbrella gives Ant Lingbo a serious edge: access to a deep pool of talent, capital, and resources. Ant has already invested in several embodied intelligence startups, including Xinghaitu, Yushu Technology, and Lingxin Qiaoshou.
But whether you look at it from a data hunger or a competitive positioning angle, Ant Lingbo can’t afford to keep its work locked up in the lab. The company has already kicked off full-scale commercial testing with ecosystem partners like Leju and Taihu, and companies like Guoda Pharmacy and Longsheng, in areas like retail sorting, logistics sorting, and industrial applications. “In the future,” Zhu Xing said, “we plan to offer our model’s capabilities commercially to a wide range of embodied intelligence customers, even to other companies building robot brains.”
At the upcoming 2026 World Artificial Intelligence Conference, Ant Lingbo will showcase the real-world capabilities of its full-stack brain 2.0. It’s worth noting that Zhu Xing repeatedly stressed in the interview that embodied intelligence is still in a very, very early stage. “The technology isn’t settled yet. From that standpoint, the opportunities and challenges are the same for big companies and startups.”