Ex-Apple AI Lead Returns to China to Tackle Embodied AI with Groundbreaking VLOA Model

Avatar 0

Reported by | Lu Keyan

Edited by | Liu Fangyuan

The debate on the future tech direction for embodied intelligence just got a fresh perspective.

At the end of June, the embodied AI firm RoboScience unveiled its general-purpose large embodied model, Visics, alongside a new technical architecture called VLOA (Vision-Language-Object-Action). During the launch event, RoboScience demonstrated the model tackling complex real-world tasks, including the notoriously tricky job of assembling furniture.

Founded at the end of 2024, RoboScience was established by Tian Ye, formerly the Head of Technical Leadership for Apple’s AI Platform, and Shao Lin, an Assistant Professor at Nanyang Technological University. The company has already secured multiple rounds of funding, most recently closing a 1 billion RMB Series A in May. To date, total funding has reached into the billions, backed by investors like 01 Capital, JD.com, China Merchants Venture Capital, SenseTime Guoxiang Capital, Puhua Capital, and DVC Capital.

Currently, the embodied AI field follows two main technical paths. The first is VLA (Vision-Language-Action), where robots mimic human movements using massive datasets of human demonstrations. The upside? They understand natural language commands directly, and the training pipeline is relatively mature. However, this approach is heavily tied to specific hardware; swap out the robot, and you often have to retrain from scratch.

The second path relies on World Models: the system learns to predict how environments and objects will change physically before acting, essentially letting the robot “rehearse” outcomes in its mind first. While theoretically offering stronger generalization, this route comes with a steep price tag in terms of training costs and significant engineering hurdles.

Enter RoboScience’s VLOA architecture. Think of it as inserting an “O”—for Object Trajectory—right between Vision-Language and Action.

Tian Ye told NUPIAO and other media outlets that the complexity of embodied AI lies in covering three dimensions of diversity simultaneously: handling various tasks, manipulating objects with different attributes, and adapting to robots with different configurations. Without a unified format to encompass all three, true generalization is nearly impossible. It’s similar to how Tokens work in Large Language Models.

In his view, the dynamic trajectory of an object serves as the “Token” for embodied AI—it tracks the position and shape changes of an object in 3D space. Unlike VLA, which is naturally bound to hardware, this approach decouples from the start. It focuses purely on how the object changes, unaffected by the robot’s body, task type, or environment, granting it inherently superior generalization capabilities.

The Visics model consists of two parts: an embodied world model that deduces object movement routes after receiving visual input and language commands, and a universal manipulation model that converts these deductions into specific instructions any robot can understand. These two are linked by continuous 3D point cloud trajectories of the objects. RoboScience’s logic is simple: data dictates the ceiling of a model’s capability, while this architectural design determines exactly what the model can learn.

Before embodied AI hits mass adoption, every single vendor faces the same nagging question: Where do we get the training data?

Wang Tao, Executive President of RoboScience, did the math. The data volume required for embodied intelligence won’t fall short of language LLMs, yet the global accumulation of real-world robot interaction data is smaller than LLM training data by a factor of 10^6 to 10^8.

Many players believe only massive amounts of real physical interaction data can train deployable robots. That’s why, over the past few years, almost everyone has been pouring money into real hardware data collection, utilizing things like material factories and motion capture equipment.

Wang pointed out that the cost per data point using current real-hardware collection methods is roughly a few cents, with each person collecting maybe a few hundred clips a day. The entire industry’s monthly capacity sits at the tens of thousands level. Especially during post-training phases, complex single-task demonstrations require tens of thousands of manually annotated data points, causing labor and time costs to scale linearly with task count. More critically, data collected in controlled factory environments suffers from distribution shifts compared to real-world scenarios, making stable generalization in actual deployment a nightmare.

RoboScience chose a completely different road.

Since pre-training requires diverse and massive datasets that are hard to find in the real world, RoboScience relies on internet videos and their proprietary simulation engine, RoboMirage, to generate high-quality data before deploying to real scenes. Real hardware data is reserved for post-training on specific scenarios, providing those difficult failure cases that simulations alone might miss.

According to Wang’s calculations, this data production flow ties only to computing power, not human labor. The cost per data point drops to mere fractions of a cent—just 1/20th to 1/200th of traditional solutions. Plus, theoretically, adding more GPUs means unlimited scalability.

Currently, RoboScience has accumulated millions of hours of video data and billions to hundreds of billions of simulated operation trajectories. This year’s goal? Surpassing 10 million hours of video data and hitting trillions of simulated data points.

At the launch event, RoboScience showed off a robot autonomously reading IKEA assembly manuals to put together furniture. Even when humans dismantled assembled parts mid-process, the robot automatically recovered and finished the job. Other demos included tying a tie, balancing a coin, opening envelopes, and grabbing potato chips or eggshells. Notably, the tie-tying task was trained entirely on simulation data.

There’s a growing industry consensus that 2026 won’t be the “ChatGPT moment” for embodied AI. More vendors are dropping the rush for full-scenario generalization, focusing instead on specific use cases to prove business viability before expanding their generalization boundaries. In a way, RoboScience took the opposite path: building a relatively universal base model first, then validating and refining it through real-world scenarios.

Tian Ye believes that while base model iteration and scenario deployment aren’t mutually exclusive, the choice of scenario dictates the future tech roadmap. Narrowing down to specific scenes often leads to overfitting with small data and small models, whereas demanding high generalization forces the base model to keep evolving.

In his eyes, since base models are the foundation for many applications, RoboScience chose to let scenarios drive model training from day one, ensuring relative generalization capabilities. Simultaneously, they are developing their own robot bodies to ensure deep coupling between hardware and scenarios.

No matter how much the base model iterates, it eventually comes down to commercialization. RoboScience currently has three main paths: licensing pure software capabilities (already generating revenue, mostly to robot manufacturers and integrators), providing domain controllers loaded with their custom large models for industrial or collaborative arms, and selling their own proprietary robot bodies to close the loop on both commerce and data.

Wang noted that for initial deployments, RoboScience will target logistics, supermarkets, and retail sectors. These industries best showcase their tech advantages over traditional non-standard automation and offer the earliest path to monetization. When asked about profitability timelines, he emphasized that costs must drop on both the model and hardware fronts; only by scaling up will large-scale profits become possible.

The next big test for RoboScience is their self-developed robot body, set to release in August. Whether the VLOA architecture can deliver the expected generalization in real-world settings will be the first major checkpoint for this entire technical strategy.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.