Tang Daosheng Meets Yao Shunyu: Decoding Tencent’s Take on the “Second Half of AI”

Avatar 0

Reporter | Wu Yangyu

Editor | Wen Shuqi

On June 5, at Tencent Cloud’s AI Industry Conference, Tencent Senior Executive Vice President Tang Daosheng and Chief AI Scientist Yao Shunyu sat down for a candid, hour-long conversation.

This marks one of Tencent’s rare public statements since diving headfirst into the AI arms race, and the company’s first comprehensive breakdown of what the “second half of AI” actually looks like.

Over the past three years, China’s large language model scene has been nothing short of a grueling sprint. From parameter counts and training compute to leaderboard rankings, the industry spent ages fixated on raw model capability. But as foundational models start to converge, a much more grounded question has emerged: once the underlying tech methodology stops being the ultimate gatekeeper, what really dictates an AI’s actual power and commercial value?

Yao calls this phase the “second half of AI.”

In his eyes, the industry’s old mission was hunting for better ways to solve problems. Today, the heavier lift is figuring out which problems are even worth solving. With pre-training and post-training tech maturing, LLMs have essentially become a plug-and-play utility—like a universal hammer that can drive any nail. When the methodology is locked in, what truly matters shifts away from the model itself and lands squarely on scenarios, product design, operating environments, and user context.

That’s exactly why he eventually decided to join Tencent.Yao believes finding the right problems is getting tougher, but Tencent sits on a goldmine of product scenarios and massive contextual data—exactly what you need to surface those “good problems.”

“The environment is everything. Without the right setup, agents can’t really get anything done.” To Yao, the AI competition is pivoting from pure capability clashes to environment wars. Models need tools, memory, and rich context to operate. Tencent, with its WeChat ecosystem, cloud infrastructure, enterprise services, content platforms, and vast user behavior trails, naturally provides the perfect petri dish for training and stress-testing agents.

Culture gets a serious spotlight too. Yao argues that building an organization truly geared for AGI requires three overlapping gears: foundational model R&D, product落地 (go-to-market execution), and frontier exploration. Bridging those three demands radical honesty and trust. In a tech cycle defined by uncertainty, low ego, long-term thinking, and a tolerance for trial-and-error will always outweigh chasing short-term KPIs.

For the past year, if you’ve scrolled through tech commentary about Tencent’s AI strategy, “slow” was probably the word du jour. Compared to startups rushing out ChatGPT clones or internet giants loudly betting the farm on base models, Tencent has played things surprisingly cool. Neither the Hunyuan model nor the Yuanbao app came out swinging aggressively in the early days.

Tang didn’t dodge the question during the chat.He acknowledges the outside world expects more, but frames AI not as a 100-meter dash but as a decade-plus marathon. Tencent’s business matrix is incredibly complex, and different tracks demand different paces. Some areas need rapid deployment; others require longer runway to figure out the right angle.

And honestly, the game is shifting now that we’re stepping into the Agent era.The last round of competition was all about parameter scales and training throughput. The next leg prioritizes scenario density, context quality, and seamless product synergy. Tencent’s unfair advantage? Over two decades of accumulating real-world usage data across social, gaming, content, and enterprise verticals.

We’re running a marathon here, so please keep pushing us, sharing your thoughts, and actually using our products to give us real, constructive feedback.” Tang put it plainly.

From Yuanbao and Workbuddy to CodeBuddy and enterprise agents, Tencent is trying to weave together an ecosystem where models, tools, products, and scenarios feed off each other. Viewed through this lens, Tencent seems less obsessed with manufacturing a viral hit app right now. Instead, they want model capabilities to continuously loop back through diverse products for real-world calibration, while those same products get refreshed interaction patterns and service layers from the evolving models. It’s a self-reinforcing flywheel.

This conversation lays out Tencent’s core thesis on what’s next in AI. For China’s AI sector, still racing through rapid evolution and fierce market reshuffling, it’s a highly watchable signal for how the second half plays out.

Tang Daosheng and Yao Shunyu (Photo: Tencent)

Below is a slightly edited transcript of their conversation:

Tang Daosheng:A warm welcome to Shunyu.

Yao Shunyu:Hello everyone. I usually hang out in Haidian District, rarely make the trek to Chaoyang, so I’m genuinely glad to be here.

Tang Daosheng:Today’s format is pretty fresh. Hopefully, it brings some surprises. Shunyu, before you joined Tencent, I remember asking you: Why pick Tencent for the “second half of AI”? What do you see as the absolute must-haves in this phase?

Yao Shunyu:Let me break down what I mean by “second half.” I coined the term in a blog post last year. Before last year, AI development was heavily focused on engineering better methods to solve known problems. Recently though, those methodologies have matured massively. Suddenly, hunting for the “right problems” became the real bottleneck.

We used to build specialized architectures for specific tasks—like AlphaGo for chess or dedicated neural nets for machine translation. But after pre-training and post-training took off, we got a “universal hammer” that can tackle almost any nail. That turned AI into a general-purpose toolkit. Now, pinpointing the right scenarios and framing the right problems is what actually separates winners from followers.

Pulling to Tencent made sense because there’s a steady stream of good problems and mature products here. Top-tier products dictate where a model actually creates user value. Plus, environment and context are non-negotiable. You can’t order takeout if the agent lacks a delivery tool. You can’t build moats if you’re blind to what users or enterprises are actually doing. Tencent dominates the raw input layer and possesses unmatched contextual depth.

But the biggest draw? Culture. When I first chatted with you and other leaders, what stuck with me was the brutal honesty. You call out wins, you call out misses, straight to the point. Tencent runs on trust, not just cold metrics. That low-ego, pragmatic, long-term mindset is absolutely critical if you want to build an org capable of tackling AGI.

I believe the core of the AI second half lies in building a balanced, long-term “triangle” organization right here in China:

Foundation:Solidify pre-training and reinforcement learning. That takes serious compute resources and, more importantly, the discipline to execute correctly.

Products:Deliver tangible value to people and society. That demands sharp product intuition.

Frontier Exploration:Push into uncharted research paradigms. Domestic frontier work is still lagging, and I want to embed that exploratory DNA deeper into the team.

Tang Daosheng:That pragmatic vibe is exactly what clients tell us too. The AI track is a marathon. We have to face our strengths and gaps honestly. It’s a multidimensional race. Models are advancing, product formats are evolving, and the outlook is genuinely bright.

You just mentioned products providing environment and context. Internally, we constantly talk about “Co-Design”—how tightly we can fuse models with our product portfolio like Yuanbao, AI Search, and CodeBuddy. How do you approach this kind of collaborative workflow?

Yao Shunyu:I’d break it down into three pillars:

First, the base model has to be bulletproof. Pre-training is fundamentally agnostic—it fuels generalization, and improvements here lift every downstream task. Post-training, however, hinges on setting up proper evaluations. There’s a trend locally of grinding leaderboards, but what actually moves the needle is building realistic evaluation frameworks grounded in real applications.

Second, practical utility trumps benchmark stacking. Deep Co-Design with product teams boils down to earning trust. Leveraging live product data, closing the feedback loop, and polishing edge cases—that takes relentless iteration and mutual buy-in.

Third, the fundamental shift in the LLM era is generalization. Used to be you only needed translation data for translation tasks. Now, even building a coding agent requires razor-sharp chat, search, and instruction-following capabilities. That complex “data taxonomy” demands serious engineering taste.

When you operate at this level, collaboration amplifies results. For instance, the chat and search capabilities we co-designed with Yuanbao can seamlessly transfer to ima or Workbuddy. Data flowing through different products generalizes across the system, triggering genuine network effects.

Tang Daosheng:External leaderboards definitely serve as one type of evaluation. So what’s the concrete difference between our internal evals and those public rankings?

Yao Shunyu:Leaderboards offer reference points, but they’re notorious for causing overfitting. Research grounded in production data delivers way more ROI: First, it exposes baseline flaws. One major reason we ship preview models is to harvest real user feedback and patch gaps that static benchmarks simply miss.

Second, it gives us a window into actual prompt distribution. Benchmark questions are usually crystal clear and perfectly structured. Real users, though, type vague two-liners and keep probing. Those messy, organic scenarios force us to rethink training strategies. Sometimes they even spark entirely new evaluation categories we hadn’t considered.

Tang Daosheng:I still remember early Yuanbao hitting walls with multi-turn consistency. Users iterating on prompts really do play out differently than textbook benchmark examples.

Yao Shunyu:You’ve fired enough questions at me—I’m turning it around. During our first chat, you walked through your journey from QZone and QQ Avatars to QQ Music, then to cloud computing and now Yuanbao. You’ve shipped to C-end and B-end, bridging ancient mobile eras and modern AI. What’s your North Star for product creation? What principles never change? What’s completely flipped?

Tang Daosheng:The core logic hasn’t budged: products must satisfy user needs, crush pain points, and deliver measurable value. Otherwise, nobody pays. Whether it’s PC, mobile internet, or industrial internet, the底层 logic stays identical.

That said, AI absolutely changed the playbook. First, the paradigm shifted. Product building used to be feature-led—like a cafeteria menu where users pick what’s available. AI opens the door to unstructured service. Interactions run on natural language, and you can’t predict every query. That forces you to lean on the model to grasp intent, then chain reasoning to trigger the right tools.

Second, the workflow transformed. We went from waterfall development with rigid specs and QA gates to something fluid. LLMs now draft code, so engineers pivot toward system architecture and continuous guidance. Testing has to shift left—meaning you bake in evaluation environments, align open-ended outputs, and calibrate tone/style much earlier. Shipping AI products now demands a completely hybrid skillset.

Yao Shunyu (Photo: Tencent)

Yao Shunyu:It’s undeniably harder.

Tang Daosheng:Yeah, it is. Everyone’s buzzing about the Hunyuan 3 Preview as your Tencent debut. What specifically shifted under the hood?

Yao Shunyu:Not really magic. Building a top-tier model is actually quite mundane—it comes down to nailing infrastructure and data.

We completely rebuilt the stack, covering both pre-training pipelines and reinforcement learning loops.

We overhauled data curation and evaluation protocols, redefined what counts as a “real” problem, and dialed up both data quality and classification granularity.

Lots of moving parts—hiring cadence, trade-offs, resource allocation—don’t yield clean formulas. They run on engineered taste and instinct.

I’m curious though—how do you view Co-Design now? Where does the model own the wheel, and where should product take the driver’s seat?

TangDaosheng:Co-Design has been in constant flux these past two years, largely riding the wave of model upgrades. My biggest takeaway? Alignment is brutally difficult. The product side wants to solve a specific user friction; how does the model adapt? How granular should labeling be? What constitutes a meaningful reward or penalty?

If the product team’s definition of a great experience clashes with the model’s evaluation metrics, you end up shipping conflicting features. Co-Design forces cross-functional alignment on open-ended targets. Miss the alignment mark, and the product behaves erratically—or worse, randomness creeps in because the training signal got noisy.

Yao Shunyu:I think trust is the actual hard part. Model researchers chase maximum capability; product managers chase user satisfaction. Those incentives naturally diverge, so empathy becomes the bridge.

Taking Yuanbau as an example: our pre-training wasn’t quite ready yet, but I dispatched our strongest post-training lead to bolster the Yuanbao squad. Some algorithm folks pushed back, but I knew protecting Yuanbao’s daily active users was critical for future model iterations. That deliberate pivot showed the product team the model group was genuinely invested in their success. That trust paid dividends later when Hunyuan 3 Preview launched smoothly on Yuanbao.

Tang Daosheng:Pivoting topics slightly. You pioneered the ReAct architecture, and your PhD revolved around Language Agents. Have some of your older theories finally come due? Which ones?

Yao Shunyu:I actually felt pretty emotional re-reading my dissertation recently. It’s from 2019, titled “Language Agents for Digital Automation: From Next Token Prediction.”

Tang Daosheng:That’s already seven years ago.

Yao Shunyu:We were living in the GPT-2 era. Outputs were choppy, riddled with artifacts. Hard to imagine it shaking up the world. Researchers back then played it safe—get a model to predict “Beijing” when asked about China’s capital, and everyone claps.

My imagination ran wilder. Even though GPT’s token-generation mechanism was brutally simple, I saw it as wildly universal. I believed it could automate literally everything digitally. Looking back, maybe it unlocks dual automation across both digital and physical realms.

During my PhD, I tackled two main streams: first, architecting the agent methodology—turning a “prediction engine” into an “automation engine.” That’s where ReAct landed. I’ll never forget a night in July 2022 when I first hooked the model up to a hand-written Wikipedia API. It successfully navigated a multi-turn web interaction for the first time. It felt like a dim bulb suddenly flipping on. I knew this would reshape the landscape within five to ten years. I just didn’t expect the timeline to compress so violently.

Stream two involved defining “digital automation” benchmarks like WebShop and SWE-bench. Looking at it now, external-facing agents and coding agents are definitely the two biggest technical branches today. My dissertation’s Future Work section listed: training models for agents, robust deployment, scientific discovery, and human assistance. I’m incredibly lucky to actually be executing on that exact roadmap today.

Tang Daosheng:Flawless execution. Those directions are rolling out across the industry. But tech often outpaces hype. Agents currently guzzle tokens. For Hunyuan’s next generation, where are you placing your bets? What matters most right now?

Yao Shunyu:Zero hesitation—agent proficiency and coding capabilities are now table stakes, sitting right alongside pre-training fundamentals. I consider the Coding Agent deeply structural. Logically, it’s Turing-complete: once a model controls the filesystem and container orchestration, it essentially becomes a self-contained system.

We’re approaching it through a few distinct lenses:

Holistic Data Mix:Nailing coding isn’t just about feeding it code. You need conversational data, reasoning traces, and everything else. The LLM’s superpower is generalization, so the training diet must reflect that.

Product Feedback Loops:How do we mine real-world online traffic for corrective signals? That demands mature Co-Design muscle.

Structured Imagination:Beyond incremental engineering gains, we have to carve out space for speculative, high-uncertainty research that hunts for the very next paradigm.

Tang Daosheng:From a product standpoint, everyone’s wrestling with “token anxiety.” Costs are skyrocketing. Clients and internal teams are watching token burn rates like hawks. Any playbook for squeezing better efficiency out of models during execution?

Yao Shunyu:When we talk cost-performance in China, eyes immediately jump to architecture. But it’s a systemic equation.

First, raw performance is the ultimate cost-saver.Ironically, many teams find deploying the beefiest model actually cuts costs. Why? Because it nails the task on the first try, eliminating repeated retry loops and manual intervention overhead.

Then there’srobustness, especially for straightforward tasks. If a leaner model can match heavyweights on routine workloads while staying rock-solid, that’s a massive win in the Chinese market.

I’m also curious—when did you realize agents represented a genuine product breakout? What’s your current mental model? Where do you see the real friction blocking wider agent adoption?

Tang Daosheng:Agent UX morphs depending on the scenario. Design’s core job is unlocking the model’s latent abilities. Surprisingly, as models iterate, the agent workflow is actually simplifying: we’re primarily handing models better toolkits (Skills), persistent memory, and user preference profiles to ground their reasoning.

So the real bottleneck is contextual filtering. We have to figure out which data slices matter for a given scenario, extract them cleanly, and sync them with the model so it has ammunition when it needs to reason.

Yao Shunyu:Workbuddy’s picking up serious steam lately, backed by rapid cycles from tiny squads. Question: compared to traditional software dev, how has the R&D rhythm and org management shifted in the Agent age?

Tang Daosheng:I’ve been auditing Workbuddy’s structure recently. It’s radically flat compared to legacy setups. Mostly pods of 3 to 5 people, laser-focused on a single domain, running constant experiments.

This org shape demands psychological safety around failure. Most experiments flop, but you need that volume of trials to isolate workflows that genuinely smooth out user journeys.

Also, role boundaries are dissolving. Engineers aren’t just typing syntax anymore (AI handles that boilerplate). They’re becoming idea-driven architects steering multiple coding agents. Simultaneously, devs have to shift-left into evaluation design and alignment tuning.

Tang Daosheng:I want to touch on a hot-button topic. Critics claim Tencent moved too slowly on AI and missed early windows. Were we actually late? What defines this second half anyway?

Yao Shunyu:Should’ve been my question to you (laughs). It ultimately boils down to two convictions:

Number one: Is AI a sprint or a marathon? Silicon Valley voices sometimes argue AI will automate every job within eighteen months, so cash out and retire. Our stance? The game just tipped onto field. The second half is barely underway. ChatGPT or Claude won’t monopolize the super-app throne forever. Fresh opportunities will keep cascading in.

Number two: Linear trajectory or multipolar ecosystem? Sure, everyone’s burning calories on pre-training and coding agents right now, making it look like a single lane. But reality will fragment. Multimodal systems and embodied intelligence are only just cracking open.

Viewed through that lens, if the second half just kicked off, panic buying makes zero sense. Taking scenic routes is normal. What actually separates players is whether they can absorb honest feedback and maintain strategic patience.

Tang Daosheng:We actively welcome higher scrutiny. Tencent operates across a sprawling business landscape, meaning pacing varies wildly per vertical. Some units accelerate; others are still mapping the terrain.

As you noted, this is a marathon. Tencent’s moat is contextual richness. AI starves without context, and Tencent’s years of cross-domain accumulation supplies the most commercially viable context for model refinement.

We actually responded to the Agent wave faster than outsiders realize. Products like Workbuddy didn’t appear overnight—they evolved organically from CodeBuddy experiments years back. We spotted massive untapped demand from non-engineers and pivoted quickly. Enterprise clients are already placing serious bets on our combined suite.

Due to time constraints, we’ll wrap here. Huge thanks to Shunyu for the transparent deep dive.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.