Journalist |
Editor | Wen Shuqi
“When we decided to push MaaS as our top priority three years ago, a lot of people in the industry thought selling tokens was a money-losing game.” Reflecting on how Volcano Engine built its market moat, President Tan Dai said that strategic foresight is one of the key ingredients for staying competitive.
IDC data shows that Volcano Engine commands a 49.5% share of China’s public cloud MaaS market. And according to the company, the Doubao foundation model’s average daily token usage had already surpassed 180 trillion by June this year, growing more than tenfold in the last twelve months.
Now, at the launch of a new flagship model generation, Volcano Engine is doubling down on its ability to execute complex tasks in real-world scenarios while keeping costs low, aiming to cement its lead in the enterprise AI space across China.
On June 23, at the 2026 FORCE Conference, Volcano Engine took the wraps off several new models, including the Doubao 2.1 Pro, covering video, image, and audio capabilities.
During the event, Tan emphasized that a large model’s capabilities must cross a “productivity inflection point” to truly meet production needs. He pointed to Claude Opus 4.6 as the first model to have done so. The star of the show, Doubao 2.1 Pro, brought upgrades in three critical dimensions: coding, agentic behavior, and vision-language understanding (VLM).

Based on industry benchmarks the company shared (like SWE-Pro and OSWorld), the new model’s coding and multimodal performance matches or even surpasses that of top-tier overseas models, including Claude Opus 4.6.
In live demos, Volcano Engine showed Doubao 2.1 Pro running a chip design RTL test continuously for nearly 18 hours and completing the entire engineering workflow – a testament to its stability on long-cycle, complex tasks. In another demo, the model coordinated over 500 agents working together to generate a 3D virtual city.
Pricing for Doubao 2.1 Pro is set at RMB 6 per million input tokens, RMB 30 per million output tokens, and as low as RMB 1.2 for cache hits. Volcano Engine claims that overall usage costs are almost 80% lower than Claude Opus 4.6. On top of that, a Turbo version optimized for high-frequency scenarios halves the price again.
Doubao 2.1 is now available via Volcano Engine’s API service and is being rolled out across products like Doubao, TRAE, and Coze.
The company also teased Seedance 2.5, a video generation model due in July that can produce 30-second native videos and is already finding its way into embodied intelligence, industrial manufacturing, and autonomous driving – real-world industries that need real-world video.
The state-of-the-art performance of the Seedance family and the surging demand for AI short dramas have become two forces that feed each other, making Seedance the productivity workhorse of that market. But Volcano Engine clearly isn’t stopping there. Tan told NUPIAO that the team still wants to see Seedance go deep into production environments across all sorts of industries.
“Especially for high-end manufacturing and building world models – that’s what we care about most,” Tan said.
On the much-discussed issue of Seedance supply, Tan stressed that video generation models are structurally different from coding and agent models. They’re mainly based on Diffusion architectures, which have relatively low demands on underlying chips – especially high-bandwidth HBM. Because of that, the team has done a ton of optimization on the Volcano Ark inference platform, allowing Seedance to tap into all kinds of compute power, including lower-end chips.
“So there’s absolutely no conflict with coding and agent models for compute resources. Resource allocation simply isn’t a problem,” he said.
In a media roundtable after the conference, Tan also talked about API pricing trends, how the Seedance video model is evolving from content generation toward world models, whether selling tokens is actually a good business, the current competitive landscape of AI coding, the next stage for large-scale agent deployment, and more.
When asked about the recent market frenzy that has pushed AI company valuations to new highs, Tan responded to whether Volcano Engine might spin off and go public on its own. “As far as I know, there are no IPO plans at the moment,” he said.

Here’s the lightly edited transcript of our interview with Tan Dai:
NUPIAO: Recently, API pricing for domestic large models has been mixed – some going up, some down. From both Volcano Engine’s own operations and the industry cost trend perspective, how should we understand the pricing strategy for your new model today?
Tan Dai: First, we can’t look at a model’s price in isolation. You have to look at price in relation to the value it creates. When a model can do more things, the value it brings is larger. Looking at both price and value, while the list price per token is rising – both for us and for other major models in the market – the value generated per token is rising even faster. So in terms of cost-effectiveness, it’s actually improving. That also ties into the fact that today’s models, whether for coding, agents, or video generation, have genuinely crossed the “productivity inflection point” and can deliver way more real-world value.
NUPIAO: Video models like Seedance 2.0 have brought incredible single-day revenue capacity to Volcano Engine. How sustainable is this growth momentum? And, since Seedance has basically conquered the short drama industry, is the demand side for video models already saturated?
Tan Dai: That’s an excellent question, and I want to make three points:
First, all the revenue figures floating around about Seedance are wrong – and inflated. I’d actually ask everyone to stop spreading them, because they put a lot of pressure on me; our finance team keeps asking me if I’m hiding revenue (laughs).
Second, short dramas are really just one piece of the Seedance puzzle. In the long run, it might turn out to be just a small niche. We’re already seeing Seedance used widely across many industries. For example, in manufacturing and retail, companies are using it to create product explainers tailored for different countries, languages, and audiences – helping them connect with consumers more effectively. We’re also seeing knowledge industries and education use it to turn knowledge into video. In high-end manufacturing, like embodied intelligence, many companies are using Seedance for data synthesis to break through the bottleneck of getting real-world data. Autonomous driving companies are using it to synthesize extreme weather or edge cases so their algorithms can become more robust.
Third, from this perspective, we believe Seedance is actually the foundation for building a “world model.” Video generation enables large-scale unsupervised learning with the fewest assumptions about data, making it a very effective way to synthesize world models. Building a good video model requires strong underlying capabilities, and it essentially leans on the strengths of the Doubao model family. Today’s Doubao 2.1 Pro surpasses Claude Opus 4.6 in coding and agent capabilities and has crossed the production-grade threshold. So the real excitement for Seedance lies in broad applications across all industries and in serving as the basis for world models.
NUPIAO: We’re hearing that Seedance is becoming a critical API procurement source for more and more video production service providers and agencies. How does Volcano Engine view this industry trend, and what kind of ecosystem role do you want to carve out? Also, from a compute security perspective, how are you handling resource allocation right now?
Tan Dai: Let me address the compute question first. Video generation models (Seedance) and coding/agent models are structurally different. Video models are mainly based on the Diffusion architecture. They have relatively low requirements on underlying chips, especially high-bandwidth HBM. We’ve done a massive amount of optimization on the Volcano Ark inference platform, making it very easy for Seedance to utilize all kinds of compute, including lower-end chips. So there’s absolutely no conflict with coding and agent models for compute resources. Resource allocation simply isn’t a problem. This is also a big reason why Seedance can be adopted at such a large scale – we’ve put tremendous innovation into both model architecture and engineering.
As for the ecosystem, we want Seedance to go deep into production environments across all industries, especially high-end manufacturing and building world models. That’s what we care about most.
NUPIAO: On the topic of “world models,” there are many different schools of thought right now. What’s ByteDance’s approach in this area? Does your training data rely mainly on video data, or on real interaction data from embodied intelligence?
Tan Dai: Internally, we’re also exploring multiple paths in parallel. But from what we see, video generation is a very effective route for synthesizing world models. It makes the fewest assumptions about existing data and can directly leverage massive amounts of video through unsupervised learning. Right now, many embodied intelligence companies are using Seedance to synthesize data and feed that back into their own model training, so we’re very bullish on this path. That’s where Seedance’s even bigger value will be in the future.
NUPIAO: When you talk about multimodal generation, you still discuss the boundary between “can generate” and “commercially usable.” How does your team define that boundary? What are the remaining bottlenecks for large-scale commercial deployment of multimodal models?
Tan Dai: The concept of the “productivity inflection point” we mentioned today is very important. Defining that boundary is actually quite simple: look at the existing business processes in each industry and figure out what model capabilities are required for each process. Once the model meets those requirements, you’ve crossed the boundary.
Data doesn’t lie. Before Seedance 2.0 came out, many people said video generation was a toy. And the data backed that up – weekend usage was far higher than weekday usage, which meant people were mostly playing around during their leisure time. But after Seedance launched, the data flipped: weekday usage is now much higher than weekend usage. This clearly shows that people are genuinely using it in real work, production, and data synthesis scenarios. That’s the productivity leap.
NUPIAO: A question on behalf of my boss – they’re a loyal Doubao user and recently felt that the quality has dropped a bit. Are you preparing a paid version?
Tan Dai: I myself use Doubao heavily every day, and I haven’t noticed any decline in quality. Also, let me clarify: the Doubao App actually falls outside Volcano Engine’s scope. But as far as I know, the Doubao App will remain free and continuously serve its broad user base with high quality – that commitment hasn’t changed. Soon, it will launch a professional task mode for productivity scenarios, powered by our newly released Doubao 2.1 Pro model. Volcano Engine’s API has been paid from day one, so there’s no such thing as degrading quality to push a paid version.
NUPIAO: Which is the higher priority for Volcano Engine right now: advancing core model capabilities, or focusing on the Harness (execution environment / toolchain) pattern?
Tan Dai: Both are critical. Volcano Engine’s ultimate mission is to help enterprises and developers solve real business problems. What customers want isn’t just a great model or a great API; they need a complete AI solution and a Harness that can land inside their enterprise environment.
That involves how the model integrates with internal enterprise systems, how it works with enterprise data, and how to handle security, agent identity authentication, and compliance requirements. That’s why we emphasize the “AI Cloud-Native architecture” today: from the bottom-level model, to the mid-layer MaaS (which includes some Harness), to the upper-layer Agent Kit (with more Harness tools), and finally to the top-layer AI workspaces. We offer zero-code, low-code, and high-code options, all designed to meet the diverse needs of different roles inside an enterprise.
The priorities evolve in a staggered way: when the model hasn’t reached the inflection point, improving the model is the most important thing. Once it crosses that point, Harness and real-world deployment become equally important.
NUPIAO: Zhipu’s valuation in Hong Kong is very high, and overseas players like Anthropic and Google have made major breakthroughs in AI coding. How is Volcano Engine positioning itself to close the gap with the most advanced models? And how do you view the market’s high expectations?
Tan Dai: Being agent-oriented is something we take very seriously. Coding is just one way a model demonstrates its abilities, but it’s extremely important because it shows the model has strong generalization power – it can auto-invoke tools and even write its own software to make up for missing tools.
Claude Opus 4.6 was the first model globally to cross the “productivity inflection point.” This year we’ve seen more models cross that threshold. Our latest flagship, Doubao 2.1 Pro, has also crossed it. Looking at the evaluation data, it consistently outperforms Claude Opus 4.6 and in some scenarios even matches higher versions. That means it’s truly ready to enter production environments that involve complex, long-running tasks.
NUPIAO: Some say large models can easily fall into a margin trap of “the more active the users, the higher the inference cost,” and some competitors say selling raw tokens isn’t a healthy business. What’s your take? What are the main metrics you use to judge whether an AI product is a healthy business?
Tan Dai: I think selling tokens is a very healthy business – I’m not sure who says it’s not.
NUPIAO: On commercialization and safety: What key moves has Volcano Engine made to connect the AI technology commercialization chain? Also, there’s been controversy around copyright issues with facial materials. Seedance 2.0 imposed verification limits on faces – what safety adjustments can we expect in future versions?
Tan Dai: Safety has always been our top priority. You saw that we spent several months polishing Seedance’s entire safety strategy before officially opening the API. That covers not just IP copyright protection for the commercial side, but also user-side facial verification and the like.
In our commercial preview, we adopted an opt-in authorization model with electronic contract revenue sharing, creating a healthy commercial loop. In the future, if we do B2B digital human avatars, we’ll also use a proper authorization verification mechanism (similar to the avatar feature in CapCut).
NUPIAO: On compute, you mentioned the self-developed DPU route. What is Volcano Engine’s thinking and next steps for self-developed DPUs and the underlying compute infrastructure? And roughly what percentage of Volcano Engine’s underlying infrastructure is “domestic compute”?
Tan Dai: Volcano Engine launched its self-developed DPU not long after it was founded. In today’s large-scale AI computing, how to better offload network, storage, virtualization, and various compute workloads to improve overall efficiency – DPUs and switches play a crucial role, and we’ve been investing deeply in self-developed R&D for these. The reason Volcano Engine has been seen as a leader in AI over the past few years is very much tied to the deep cultivation of our underlying infrastructure.
As for domestic compute, we use it extensively. For example, Volcano Ark has done a ton of compute adaptation and optimization, allowing models like Seedance to make great use of all kinds of domestic and overseas compute. I can’t recall the exact percentage off the top of my head, but it’s quite substantial.
NUPIAO: In the first half of the year, competitors made a lot of noise in AI coding, while Volcano Engine seemed to be shouting louder about the Seedance video model. How do you view this competitive difference? What’s your growth expectation for AI coding in the second half?
Tan Dai: Actually, we’ve always placed enormous importance on coding. Last year at this same venue, Ding Kun gave a keynote heavily focused on coding, back when many competitors hadn’t even started pushing it. In the first half of this year, it might have seemed like Seedance was getting all the attention, mainly because it was genuinely the global SOTA at the time, which naturally drew a lot of interest.
But internally, we’ve always believed coding is the more core and more important capability. In the second half of the year, we’ll be doing a lot more in this area. We’re already working deeply with a large number of high-end semiconductor companies, internet firms, and SaaS companies, embedding Doubao’s code models and TRAE (our AI IDE) into their R&D workflows.
NUPIAO: Many people think Volcano Engine has a video model advantage because you can link up with the internal “Hongguo Short Drama” platform. As Seedance moves into other industries in the second half, what’s your commercial playbook?
Tan Dai: I don’t see linking up with Hongguo as an advantage. Hongguo’s strategy is completely independent. In fact, videos generated by Seedance often get rejected by Hongguo’s review process (laughs). So that’s not an advantage. Our real advantage is simply that our model capabilities are strong. To break into more industries, the core play is still to keep making the model even stronger.
NUPIAO: Many tech giants are building full-stack layouts that include AI chips. How urgent do you see this kind of layout? Does ByteDance have plans to fill the chip gap?
Tan Dai: From a cloud vendor’s perspective, I don’t think having your own in-house chip is all that important. Customers are buying your model capability; they care about whether you can solve their problems, not whose chip you’re using underneath. Look at Anthropic – they don’t have their own chips, but that hasn’t stopped them from building incredibly powerful models.
NUPIAO: Earlier this year, people started talking about “token value” metrics – like core system integration rates and automation efficiency. How does Volcano Engine push its teams to increase token value?
Tan Dai: A more capable model will definitely generate more value. But to land that value in real industries, you have to deeply understand the industry and co-create with customers. For example, we understand writing code and internet applications, but we might not understand pharmaceuticals or education.
So this year we specifically set up an FDE (Field Deployment Engineer) team that goes deep into each industry and co-creates with lighthouse customers. This way we get to better understand what AI can do for that industry, and customers get to understand AI’s potential, so we can deliver a more complete solution and let tokens truly enter live production and create value.
NUPIAO: What’s the size and background of the FDE team? Which industries are they covering so far?
Tan Dai: FDEs aren’t sales or pre-sales; they must have very strong technical deployment skills, especially in AI code deployment. We also place a lot of emphasis on having diverse industry backgrounds. For instance, someone with a bioengineering background goes to work with the biopharma industry – they bring irreplaceable know-how when it comes to real-world deployment. Right now they cover quite a few industries; key sectors like automotive, healthcare, education, finance, and semiconductors all have dedicated teams.
NUPIAO: You mentioned the Agent Kit upgrade today, including multiple tools from zero-code to low-code to high-code (like TRAE, ArkClaw, Coze, etc.). How do you think about this portfolio layout and how to cover different user groups?
Tan Dai: Inside an enterprise there are different types of people – professional developers, product managers, and functional staff like HR and finance. Their needs for an AI workspace are completely different. Some need zero-code, out-of-the-box solutions; some need low-code drag-and-drop; and professionals need high-code environments.
Workloads also vary: some are pure code development, others are general office tasks like Office/PPT work. It’s hard to have one product that rules them all right now, so we’re launching a diverse toolbox matrix that covers different dimensions from zero-code to high-code. In the future this matrix may evolve and converge, and we’ll keep iterating as AI advances.
NUPIAO: On visual models, are you leaning more toward a pixel-centric video generation route, or a route that combines multiple modalities?
Tan Dai: We’re definitely exploring multiple directions. Right now the pixel-generation route (Diffusion) is progressing faster and delivering better results, so we’re putting a bit more energy into it. But we’re also working on other streams like 3D generation (for example, Doubao’s 3D model). No matter which route, the core is always to better connect “generation” and “understanding.”
NUPIAO: Can you share a bit about Volcano Engine’s global expansion thinking?
Tan Dai: We take the overseas market very seriously. Of course, Volcano Engine itself mainly focuses on the China market; for overseas, we have another entity handling that.
But the model itself is naturally global. If the capability is good enough, it will naturally attract global customers. For example, right now almost half of Seedance usage comes from overseas, with many large multinational companies and creator platforms (like Canva) using it. What overseas users really care about is model capability and cost-effectiveness. We’ve also set up MaaS access points globally – in places like Southeast Asia, the Middle East, and Europe – to make it easier for global developers to call our services.
NUPIAO: When you internally decide whether a new scenario is worth building into a standalone Agent product or just adding as a skill module inside an existing agent, what metrics do you look at?
Tan Dai: First, we look at the commercial outlook. If the market potential for that scenario isn’t even at the 1-billion-yuan level, you’re probably better off not making it a standalone agent product – just write it up as a skill and add it in. As model capabilities get stronger, tasks that used to require complex agents can now often be solved with a skill or a dynamic workflow configuration.
NUPIAO: Looking at the current industry demand, what stage of development do you think China’s large model market is in?
Tan Dai: It’s still at a very early stage. You could say that last year we ran maybe 500 meters, and this year we’ve run a little over a kilometer. But that “one kilometer” is extremely critical, because it marks the point where model capabilities have crossed the “productivity inflection point.” Once domestic models meet or exceed that standard, it means they can actually be used in production and generate real commercial value.
NUPIAO: The conference mentioned that the Doubao foundation model’s average daily token usage has reached 180 trillion. What’s the split between internal business and external customers? And how does the compute ledger look?
Tan Dai: That 180 trillion is the total across all Doubao models – including internal business, external API calls, and the Doubao App. In pure token volume, the Doubao App accounts for a relatively large share. But if you look at economic value or the cost-per-token spent, external customers (enterprise scenarios) take a higher share. That’s because external calls are mostly used in complex, production-grade scenarios, where each token processed generates much higher value. We haven’t done an internal, consolidated assessment of the overall compute ledger.
NUPIAO: Volcano Engine now holds a 49.5% share of the domestic large model market. With such fierce competition, what moat do you rely on to defend that position?
Tan Dai: It boils down to two things:
First, model capability – especially being the first to cross the “production-grade inflection point.”
Second, how you bring the model into the enterprise. That includes the FDE deployment model, deep industry understanding, deep collaboration with ecosystem partners, and our team’s own expertise in AI solutions.
Maybe there’s also strategic foresight. Three years ago when we decided to make MaaS our top priority, many in the industry still thought selling tokens was a money-losing business. Having conviction about the future is also a key part of staying competitive.
NUPIAO: You mentioned earlier that digital employees have token usage evaluations. Do real human employees also get evaluated on their token usage? Will there be any token subsidy policies?
Tan Dai: Internally, employees can voluntarily turn on a system to track their daily token usage, but it’s purely for personal observation and efficiency improvement – it will absolutely never be used as a KPI. In practice, we’ve found that blindly applying AI sometimes doesn’t help. If the goal-setting or the first principles of a task are wrong to begin with, no amount of AI will fix it. So before using AI, you still need to go back to the business fundamentals and clarify your objectives.
NUPIAO: How do you view the recent phenomenon of AI company valuations hitting new highs? Does Volcano Engine have any plans to spin off and go public on its own?
Tan Dai: As far as I know, there are no IPO plans at the moment.