Reported by NUPIAO |
Edited by NUPIAO | Wen Shuqi
“We’ve moved past purely linguistic model research. Over the last two years, our focus was squarely on multimodality, but the real future lies with world foundation models.” Says Wang Zhongyuan, Dean of the Beijing Academy of Artificial Intelligence (BAAI). According to him, AI is finally stepping into the physical world—and this marks a genuine turning point for the entire industry.
Nestled near Wudaokou and surrounded by several top-tier universities in Haidian District, BAAI has quietly become one of China’s most influential research hubs since large AI models exploded onto the scene. Outside of startup labs, you won’t find a more representative center for cutting-edge R&D.
The institute keeps churning out major milestones—from the WuDao language series back in 2021 to this year’s WuJie multimodal lineup, like Emu3.5. It’s also been an incubator for heavy-hitting startups like Zhipu AI, Moonshot AI, ModelBest, and Galactica. Some have already gone public, and their market caps reflect the massive confidence investors have in the tech coming out of here.
But BAAI isn’t resting on its laurels. At the 2026 BAAI Conference, they officially unveiled two fresh breakthroughs in the world model space: WuJie Physis v0.1 and WuJie RoboBrain Orca.
Physis v0.1 is built as a universal world foundation model. Its main goal? To lay the groundwork for how AI actually perceives and reasons about the physical realm. We’ve all seen earlier niche world models struggle—they often ignore real-world physics, produce shaky predictions, and lack long-term temporal memory. Physis aims to fix that by unifying perception, interaction, and decision-making across every scenario, making AI’s physical reasoning far more reliable and grounded.

Then there’s RoboBrain Orca, designed specifically as an embodied AI brain centered around predicting the next physical state. It packs three core strengths: unified representation, causal reasoning, and cross-modal decoding. By simultaneously generating language-based thoughts, visual forecasts, and motor decisions, it empowers robots to seamlessly navigate the full loop from cognition to action. Think long-term autonomous operations in real-world settings like logistics hubs or hotel service environments.
“AI is going through a massive paradigm shift—from living in the digital sphere to operating in the physical world,” notes Wang Zhongyuan. “The core driver isn’t just predicting the next token anymore. It’s about predicting the next physical state.”
Back in 2023, Yann LeCun made headlines at a BAAI conference by pointing out that language models alone will never get us to AGI. That’s when “world models” really started gaining serious traction across the industry.
BAAI doubled down the following year, declaring world models a non-negotiable path to AGI. Over the next two years, they funneled resources into a clear R&D pipeline anchored by native multimodal models like WuJie Emu3 and 3.5, culminating in this year’s strategic rollout of Physis and RoboBrain Orca.
Now, don’t confuse this with early hype around video generators. While some initially thought tools like Sora were the blueprint for world models, BAAI draws a hard line. A true world model has to actually understand and simulate real physics. It needs causal tracing, long-sequence consistency, and the ability to serve as the foundational bedrock for Physical AGI.
So BAAI mapped out a clear taxonomy for where things stand today. They’re currently blending approaches from Route 1 and Route 4, carving out what they call a potential “Route 5”:
Type 1: Language-Centric World Models, including VLMs and VLAs. These models predict the next word in text space. They learn a world described by language, but they completely miss the underlying physical consequences;
Type 2: Pixel-Centric World Models, like Sora and Seedance. These learn videos or images in visual space, essentially mastering a world reduced to pixels;
Type 3: 3D Structure-Centric World Models, covering everything from basic 3D reconstruction to Fei-Fei Li’s team’s World Labs Marble. But let’s be clear: rebuilding a 3D space isn’t the same as understanding the world, and geometric layouts don’t automatically equate to physical states;
Type 4: Visual Representation-Centric World Models, such as Yann LeCun’s JEPA series. These predict compressed visual representations, but evolving visual embeddings doesn’t mean you’ve captured actual physical laws.
When it comes to the US-China landscape, Wang thinks the playing field is surprisingly level right now. We’re all at the starting line. In this phase of figuring out physical laws and building foundational layers, Chinese research institutes are already rolling out genuinely original technical approaches.
On the elephant in the room—data scarcity—Wang remains pretty optimistic. He points out that video holds massive untapped potential. Humans pick up basic physics just by watching others move around, so video is still king for scaling up. On top of that, BAAI is actively pairing academic collaborations with real-world data collection to get ready for whatever comes next.
So where does this all lead? Wang sees embodied intelligence as the flagship application, but the real value stretches way beyond robots. From simulating protein folding at the microscopic level to optimizing macro-scale industrial manufacturing, anything that relies on modeling and predicting physical behavior falls squarely into the world model playground.
He compares where we are today with world models to where deep learning sat back in 2012. Look at the timeline from then until ChatGPT blew up the world—it took the whole decade.
This time around, though, we probably won’t wait another ten years. Wang figures that as we dig deeper into video datasets and refine physical simulation tech, the ramp-up could compress dramatically down to just three or five years.

Below is an edited transcript of our conversation with Wang Zhongyuan:NUPIAO conducted light editing for clarity:
World Models Are Fast-TrackingAIInto the Physical World—A Historic Turning Point
NUPIAO: When did the BAAI world model team actually come together, and what’s the background history?
Wang Zhongyuan: BAAI’s AI strategy has always followed a steady, pre-planned roadmap. Today’s world model team is basically a merger of two groups: one evolved naturally from our multimodal and embodied AI teams over the last couple of years; the other is the new foundation model squad we officially kicked off early this year, once the multimodal architecture path was fully validated.
On top of the existing crew, we brought in fresh talent like Chen Boyuan and Wang Pengwei early this year. Chen’s only 22, but he’s already leading BAAI’s Behavior World Model Center, smashing previous age records for the WuJie Emu team lead. BAAI has always pushed young researchers to take the helm—we care about capability, not titles or seniority.
NUPIAO: How big is the world model team right now?
Wang Zhongyuan: Not especially huge, because we’re laser-focused on R&D right now. We’ve got solid historical data, a mature engineering backbone, and some standout researchers joining us. Resources for the lab phase are plenty. We’ll only scale up significantly when we start shipping actual products like robots or running physical experiments.
We wrapped up our foundational LLM exploration two years ago and handed those torches to commercial partners like Zhipu, Moonshot, and ModelBest. Right now, BAAI’s full attention is locked on world models. We’re running alongside international peers from the same starting line. Our calls aren’t perfect, but we respect how science works and are willing to eat the risk of failure—that’s the responsibility of a proper research institute.
NUPIAOYou guys were talking about world models three years ago. If you had gone “All In” back then, would you have scaled faster?
Wang Zhongyuan: Everything moves at its own pace. Without getting our hands dirty on multimodal world models and actually mapping out the multimodal scaling laws, jumping straight to full world models would’ve been premature.
We stuck to our own playbook: mining massive datasets, fusing multimodal signals, then dabbling in embodied AI. Once we realized embodied setups couldn’t crack generalization on their own, and saw the massive open space in AI For Life Science, doubling down on world models felt like the obvious next step.
NUPIAOHow do you actually lock in these R&D directions and decide where to bet?
Wang Zhongyuan: We spend at least half our yearly calendar on internal deep-dives where the whole research staff debates AI’s trajectory. We also lean heavily on feedback from experts at the BAAI Conference. You have to build your own worldview instead of getting swept up by external hype cycles.
As for the betting logic? We believe the human brain can decode both language and physical actions. If we can build a unified representation space that branches into different outputs, the scaling potential is enormous. That’s our baseline technical conviction.
NUPIAODid you guys fight it out internally before settling on a direction?
Wang Zhongyuan: Absolutely. But as a non-profit research hub, open source and transparency are baked into our DNA. We’re happy to share unfinished ideas to spark healthy debate across the industry.
NUPIAOIs the world model initiative your absolute highest priority right now?
Wang Zhongyuan: Basically everything we’re doing falls under the broader world model umbrella, spanning both macro and micro scales. We’ve officially stepped away from pure LLM research. Multimodal world models carried us through the last two years, but the future belongs to world foundation models. Even AI For Life Science—protein folding, neuromorphic computing, you name it—is fundamentally part of this push. Eventually, you realize everyone’s just trying to model the physical world differently.
NUPIAOTackling world models right now feels incredibly steep. What’s BAAI’s attitude heading into this?
Wang Zhongyuan: We’ve always positioned ourselves as trendsetters. Whether it’s LLMs, multimodal models, embodied AI, or world models—we’ve maintained strong confidence throughout.
Our exact route hasn’t fully hardened yet, but we’ve placed our chips on latent space. We’re experimenting with compressing world knowledge into a latent layer, then using different decoders to forecast actions and states. Could be right, could be wrong, but the results in a couple of years will tell.
NUPIAOYou mentioned Physical AGI has a massive ceiling. What exactly do you mean by that? Does this year’s BAAI Conference aim to draw a technical roadmap or establish a value coordinate system?
Wang Zhongyuan: The ceiling stems from the sheer complexity of the physical world—time, space, fundamental laws, plus all the tools humans have invented. LLMs created insane value in the digital realm for writing and coding, but they still choke on physical problems. The physical world is where humans actually live and produce, and the economic upside + problem difficulty dwarf anything in the digital space.
We dropped the “WuJie” series last year, one of the first in the industry to explicitly talk about bridging the digital-to-physical gap.
For this year’s conference, we want to do two things: hash out technical paths, and firmly mark the moment AI crosses into the physical world as a historic inflection point.
NUPIAOGPT exploded in 2023. Where exactly do world models sit on that timeline right now?
Wang Zhongyuan: I keep seeing current world models and embodied AI as sitting right where deep learning did around 2012. Neural nets had depth, but they were still stuck solving narrow tasks in isolated scenarios. It wasn’t until Transformer architecture matured in 2018, followed by ChatGPT late 2022, that we hit the wall. Took a full decade.
Evolution is accelerating now. Data accumulation might hit critical mass in just three to five years. Video potential is barely tapped, and embodied robots are gathering real human-interaction data as they deploy. All of this turbocharges the world model explosion.
Deployment usually follows tech readiness. Deep learning concepts surfaced back in 2006 but didn’t truly explode until years later. We’re testing every viable path now so we’re ready when the tipping point arrives.
“VLA is the now, world models are the next”
NUPIAO: Last year everyone was hyping multimodal fusion. Now world models are the new wave. What’s the actual difference between the two?
Wang Zhongyuan: Early multimodal models (like WuJie Emu) mostly stitched together text, images, and video. Sound and action barely made the cut. Actually entering the physical world demands explicit handling of State and Action—that’s a much stricter physical constraint.
A lot of industries casually slap “world model” on video generators, but those can’t solve real physical problems. Sure, a generator can spit out a pig flying, but that violates gravity. If you wire that confusion into a robot’s brain, it might seriously misjudge itself as Iron Man and cause actual damage.
BAAI’s world models are built explicitly for real-world physics. It’s an extension of multimodality, not a replacement.
NUPIAO: There’s a ton of流派 floating around now—spatial intelligence, JEPA, diffusion models. How does BAAI’s approach actually differ from the mainstream domestic and international routes?
Wang Zhongyuan: We see four dominant technical tracks currently:
Type 1: Language-Centric World Models, including VLMs and VLAs. These models predict the next word in text space. They learn a world described by language, but they completely miss the underlying physical consequences;
Type 2: Pixel-Centric World Models, like Sora and Seedance. These learn videos or images in visual space, essentially mastering a world reduced to pixels;
Type 3: 3D Structure-Centric World Models, covering everything from basic 3D reconstruction to Fei-Fei Li’s team’s World Labs Marble. But let’s be clear: rebuilding a 3D space isn’t the same as understanding the world, and geometric layouts don’t automatically equate to physical states;
Type 4: Visual Representation-Centric World Models, such as Yann LeCun’s JEPA series. These predict compressed visual representations, but evolving visual embeddings doesn’t mean you’ve captured actual physical laws.
BAAI leans hardest into Route 4, but we’re actively fusing it with language modeling to forge a potential fifth track. I came up through vision, so I know signal quality matters, but LLMs are undeniable powerhouses for reasoning and planning. A world model shouldn’t just be a fancy simulator—it needs to be a decision-support engine for humans. Our WuJie Emu3.5 already blends multimodal generation with world-model reasoning.

NUPIAO: World models are still super early. What’s the biggest technical hurdle you need to break through?
Wang Zhongyuan: First, baking physical laws into multimodal fusion. Take a water bottle about to tip over—the cap being on or off completely changes how it falls. Humans instinctively predict that. Teaching that to a model? Still a grind.
Second, long-sequence consistency. Current video generators can stretch out clips, but they routinely break physics over time. Camera pans away, comes back, and suddenly the clock on the wall jumped three hours forward. That kind of hallucination breaks immersion and utility.
Third, action integration. Embodied AI and hardware are dumping tons of real-world data now, but it’s nowhere near enough. Just like LLMs needed the entire internet to bootstrap, world models need massive volumes of high-fidelity physical data before they truly ignite.
NUPIAOWhere will the real competition boil down to? What factors ultimately decide the winner?
Wang Zhongyuan: Everyone’s slapping “world model” on their PR decks right now, but most are just narrow tools or specific-scenario fixes. That’s not what we’re building—a universal world foundation model. We haven’t even agreed on a unified definition yet, which is why carving our own lane feels safe. You can’t compare apples to oranges when nobody’s defined the fruit basket.
NUPIAOWhat needs to click before this whole space starts converging?
Wang Zhongyuan: We need a system or product that proves it actually passes physical validation tests, maintains long-term sequence logic, and demonstrates causal reasoning. And crucially, it has to function as a base model that you can fine-tune across wildly different domains without breaking.
NUPIAOThere’s constant debate about world models versus VLA (Vision-Language-Action). Is the world model the only path for embodied AI, or can they work together?
Wang Zhongyuan: VLA is the now. World models are the next.
VLA gets robots moving fast in controlled spots like warehouse sorting. But it’s heavy, suffers from latency, and struggles with spatial generalization and complex physical reasoning. Ten years from now, we’ll likely run smoother app-level models, but if you want to handle long-horizon tasks and actually grasp physics, the world model is the necessary gateway.
NUPIAOPlenty of video model companies are pivoting their messaging to claim they’re doing “world models.” The terminology feels vague lately. How do you view this shift?
Wang Zhongyuan: Honestly? Good thing. Industry consensus means talent, capital, and engineering bandwidth flood in, which objectively speeds up progress. Yeah, we’ve got at least four rival camps right now, everyone benchmarking against each other, causing some noise—but that’s normal developmental friction.
Look at how LLMs evolved. Until the dominant path solidified, everyone pitched their own flavor. We know exactly what we want: a generalized foundation model that solves diverse downstream tasks, not just a fancy clip generator.
From first principles, humans don’t render ultra-high-res mental movie scenes to plan ahead. We just simulate probable outcomes in our heads. That’s the target.
NUPIAOCan pure video generation bypass physical interaction and spontaneously develop causal reasoning?
Wang Zhongyuan: Logic-wise, LLMs run on “Next Token Prediction,” while world models flip to “Next Physical State Prediction.” That “state” bundles language, motion, spacetime, and cross-modal data together. Pure VLMs fall short because they lack Action, and even audio signals remain poorly integrated.
Most current embodied systems are just passive command executors. We believe AI entering the physical world must possess independent reasoning and planning capabilities—able to drive agents, execute moves, and self-evaluate. The ceiling is astronomical, and the bugs to squash are endless.
NUPIAOLLMs haven’t truly cracked human cognition yet; they’re just statistically predicting text. Do you think future world models will actually grasp underlying laws, or will they just keep predicting?
Wang Zhongyuan: World models predict the “next physical state.” That state bundles text, imagery, audio, and motor outputs—far richer than pure language modeling.
Will that predictive framework birth human-like intelligence? I believe it will evolve toward it. Whether we call it “Physical AGI” eventually? People will argue semantics forever. BAAI’s job isn’t to win definition wars. It’s to ship usable capability that makes society better.
NUPIAOWith world models taking shape, can existing foundation model vendors jump in? Where does the new moat actually lie?
Wang Zhongyuan: Never rule out big tech and internet giants entering the space—car manufacturers are already circling. The industry collectively knows this is the inevitable direction.
History shows every new era spawns groundbreaking new companies. Big players are pushing massive models, but specialized foundation vendors like Zhipu still managed to carve out massive niches.
That said, LLMs already have closed-loop monetization. Commercial firms survive on profit targets, so they’re less likely to burn cash exploring dead ends like BAAI does. Someone has to take the risk and pioneer the unknown.
That’s the beauty of research: we might blaze a trail, or we might look back in two years and realize we took a wrong turn. Both are valuable.
NUPIAOThe LLM gap between China and the US was estimated at six to twelve months. Where do world models stand?
Wang Zhongyuan: I’d say zero gap. We’re sharing the same starting blocks. Global frontier research requires top minds and accumulated intuition. We’re confident we’re co-leading the next AI epoch.
NUPIAOEarly WuDao models definitely lagged behind overseas counterparts. We’ve spent years catching up. World models carry weight comparable to LLMs. Given China’s current resources and footing, how’s this race looking?
Wang Zhongyuan: We’re definitely aiming to lead. Past few years, we played catch-up on LLMs and AI coding. But once multimodal hit, BAAI started publishing independent, original technical blueprints that got international nod.
On world models, we’ve established our own definitions and technical convictions. It proves China isn’t just following along anymore—we’re actively attempting to set the pace in cutting-edge AI.
“Data is scarce, but it won’t stall iteration”
NUPIAO: Data is the obvious bottleneck. Which slice matters most, and what’s the ideal training ratio?
Wang Zhongyuan: Long term, real-world data stays fragmented and undersupplied. But drilling down to first principles, video remains the easiest dataset to scale up and the least mined resource left.
Hear me out: a two-year-old kid watches a short video of someone eating candy and immediately learns how to unwrap wrappers and thread blueberries. Video massively accelerates how human brains form world models. So yeah, video data stays absolutely critical.
We’re also feeding our WuJie Physis model with heavy real-world physical datasets and heterogeneous sensor streams. The endgame is fixing embodied AI’s generalization blind spots and giving it actual self-reasoning capabilities.

NUPIAO: Real physical data boundaries are insanely broad. If you’re hunting for it, what’s your entry wedge?
Wang Zhongyuan: Fair question. We’re trialing multiple channels: partnering with CAS institutes for ground-truth datasets, debating whether to build lightweight custom capture rigs internally. As AI hardware普及s, organic data growth will follow. These are open frontier questions we’re actively tackling.
NUPIAOWhat else makes physical data collection so painful right now?
Wang Zhongyuan: Real environments are messy. You’re syncing visual feeds, hand gestures, ambient audio, motion trajectories, plus long contextual memory windows. Capturing that cleanly costs a fortune. Currently, we hire field crews to embed themselves in actual hotels and homes, recording via portable gear.
We want world models to eventually spout emergent generalization. Not every skill comes from memorizing scraped data. Through heavy training, models should learn to logically evolve and reason about physical reality. That’s how they’ll handle edge cases they’ve never seen.
NUPIAOHow do you gauge data quality and dimensionality? Huge impact on performance?
Wang Zhongyuan: Massive impact. We’re deeply data-driven right now. Quality and mix ratios dictate model capability straight out the gate.
Is there a magic formula for good data? Nope. It’s mostly veteran researcher intuition and empirical tweaking—which is exactly the competitive moat. Validation is simple though: drop it into a robot and see if it generalizes outside training bounds. Run it through scientific benchmarks and check if the reasoning actually holds up.
NUPIAOSince real physical data is thin, can we lean harder on synthetic/simulated data as a patch?
Wang Zhongyuan: The sim-vs-real-data debate runs deep in this space. Simulated data is engineered by humans, so inherent precision gaps exist. Can you train a beast on flawed inputs? I’m skeptical.
Sims definitely fill gaps, but they’re best treated as a mixing ingredient, not the main course. Future pipelines will blend web data, simulation outputs, and scientific datasets to jointly train world foundations.
NUPIAOIf data gaps linger, world models might only work locally. Will that cap their real-world deployment and ROI?
Wang Zhongyuan: We still feel the pinch on data, but it won’t freeze overall iteration. Video’s scalability potential remains largely untouched. Core deployment zones will stay embodied AI and physical simulation engines.
Embodied AI is still grinding through factory sorting and similar niche tasks, but that gradual rollout compounds useful data. We can’t wait for a perfect dataset to start mapping paths. World models currently look like the most viable lever to break embodied AI’s core bottlenecks.
NUPIAOTo what extent can existing LLM infrastructure stack up for world model training?
Wang Zhongyuan: My take? Nearly everything carries over.
Last year’s WuJie Emu3.5 deliberately mirrored standard LLM architectures to prove scalability. Training frameworks, data toolchains, compute clusters—all highly reusable.
Handling Action and State capture brings fresh headaches, but physically, audio, vision, and motion tracking pipelines are already battle-tested in embodied circles. I’m highly optimistic about infra reuse.
NUPIAOIs raw compute still the ultimate bottleneck for world models?
Wang Zhongyuan: Compute matters, but demand hinges entirely on architectural choices.
Take WuJie Physis—it drops language systems entirely, chasing extreme compression. Compute overhead stays manageable. Meanwhile, chasing massive LLM-scale or brute-force video generation paths burns through GPUs. Routes haven’t converged yet, but rising compute floors inevitably accelerate all world model variants.
NUPIAOWill future world models improve mainly through Scaling Law, or do we still need genius-level algorithmic breakthroughs?
Wang Zhongyuan: We need both. AI history repeatedly proved Scaling Law’s muscle: from handful of transistor parameters in the 50s, to BP algorithm hundreds in the 80s, to millions post-deep learning in 2006, cruising to billions/trillions today. That march always rode parallel waves of bigger data, sharper algorithms, and stronger architectures.
If GPU throughput climbs and multimodal datasets expand, world model generalization will absolutely tighten. Of course, we’re hungry for cheaper, leaner solutions too. Human brains run on 10-20 watts and broccoli, yet deliver insane cognition. Efficiency gains are definitely out there.
BAAI’s planting seeds in neuromorphic computing and AI For Life Science, borrowing neural structures to design tighter networks. Still early innings, but promising.
“Embodied AI is currently the biggest use-case”
NUPIAOYour materials mention WuJie models cover 50 scenarios. Why lock onto those specifically?
Wang Zhongyuan: Don’t over-index on the exact number. Those fifty scenarios just prove foundation models can plug into diverse downstream tasks. That versatility is the whole point of building a base layer.
NUPIAOHow do you measure a model’s grasp of the physical world? Is there a graduation metric?
Wang Zhongyuan: Complex and long-horizon reasoning are tough to standardize. We emphasize them because current physical apps lack generalization. Step outside tight time windows, and hallucinations + logical errors spike. World models must leverage cross-modal strength to keep spatial-physical reasoning intact over extended sequences.
NUPIAOGames and metaverse demos let you drop in a photo and spawn explorable worlds. Will world models head toward that vibe instead of traditional Unreal Engine workflows?
Wang Zhongyuan: You’re describing one category in our taxonomy: 3D world generation. That tech thrives in virtual spaces, gaming, and metaverses. Valuable? Sure. But not BAAI’s current north star.
WuJie Physis targets physical simulation instead. Tools like Unreal Engine rely on hand-curated Newtonian formulas. They look photorealistic, but trained eyes spot the artifice. Human-derived equations are inherently imperfect, capping simulation fidelity.
We want data-driven models. Hit them with enough real-world volume, and the simulated physics will surpass human-engineered engines. Still theoretical today, but if product UX beats current simulators in a few years, the industry will vote with its feet.
NUPIAOSo you’re saying world models will eventually reverse-engineer undiscovered physical laws?
Wang Zhongyuan: Theoretically, yes. Just like LLMs assist modern scientific discovery today—albeit processing digital text and formulas—future world foundations will operate at a higher ceiling, holding real potential for law discovery.
NUPIAOBeyond embodied AI, where are the next massive opportunity pockets?
Wang Zhongyuan: Embodied AI sparked the mission, but Scientific Intelligence (micro-evolutionary processes) carries equal weight.
Every sector will slap “world model” on their branding soon. Our goal stays grounded: build a foundation for real physical AI that senses, understands, reasons, and decides better. Applications inevitably cycle back to reality—healthcare, manufacturing, logistics, factories. Because current models fail at physical complexity, we need this base.
NUPIAOReports claim world models slash data acquisition costs and cut R&D cycles by 70%. Your read?
Wang Zhongyuan: Plenty think world models are just fancy data synthesizers. We acknowledge video generation boosts autonomy and embodied datasets, but that’s a side benefit, not the headline act.
Core value lives in state-driven planning and decision-making. Imagine Doctor Strange previewing branching futures and picking the optimal play in real-time. That’s the utility curve.
NUPIAOMust the endgame validate world models exclusively through embodied hardware? Can a true world model exist detached from robots?
Wang Zhongyuan: BAAI’s world foundation pushes are physics-first. An ideal base model solves embodied hurdles but equally powers autonomous driving, industrial simulation, and lab science.
Embodied AI is undeniably the biggest beachhead right now. Most current bots lack physical common sense and generalization—exactly the gap we’re targeting.
NUPIAOSo what’s the ultimate finish line?
Wang Zhongyuan: Industrial deployment. Tangible societal impact. Zero paper-chasing.
Differentiating factor vs academia: we demand visible value. Open-sourcing freely is one vector. We’ve pushed 200+ models publicly over two years, racking up over a billion cumulative global downloads. That’s measurable industry contribution.
If internal teams identify spin-outs that need tighter commercial loops, we’ll incubate accordingly.