DeepSeek Funding Meeting Recap: Liang Wenfeng Breaks Down Large Model Tech and Business Philosophy

Avatar 0

Journalist | Wu Yuyang

Editor | Wen Shuqi

A month ago, DeepSeek closed its first round of external funding, raising over 50 billion RMB (about 7.4 billion USD). The post-money valuation shot past 50 billion USD (roughly 338 billion RMB).

In this round, founder Liang Wenfeng put in about 20 billion RMB himself, making him the single biggest investor. Tencent chipped in around 10 billion. CATL and its investment arm Puquan Capital together added about 5 billion. NetEase, JD.com, Monolith, and IDG Capital each kicked in roughly 3 billion. Zhengxingu Investment and Shixiang Tech each contributed about 1.5 billion.

This instantly became the talk of the town in AI’s primary market. Industry insiders were buzzing about why a top-tier model maker, once so distant from capital markets, decided to finally take investors’ olive branch. But what’s even more intriguing is the real meat behind DeepSeek’s tech roadmap and business philosophy, as unpacked by Liang himself.

Recently, a leaked transcript from DeepSeek’s funding meetings surfaced. It’s a no-holds-barred look at how Liang sees the company, the tech, and even the industry’s future. The document is packed with core takes on company culture, tech choices, business games, and computing infrastructure—enough to help you understand the counterintuitive but maybe brilliant restraint that’s defined the firm.

Below is a curated summary of that meeting transcript:

On Talent, Organization, and the Company’s Vision

1. When we first started this company, none of us were thinking, “How much money will I make?” or “Will we go public?” The first few dozen people never had that in mind. If they did, they wouldn’t have joined.

Honestly, we came into this with a lot of goodwill toward the world. We felt this was something useful for humanity, something beyond just cash.

2. Running a big company isn’t about rules and regulations. It’s about vision. And vision isn’t a slogan on the wall—it’s how you actually operate, not what you say.

So, we don’t really have an organization. We’re driven by a vision, organized around it. We don’t go around saying, “I need to hit this KPI” or “How do I assess that.” We just have the vision.

3. AGI is that bigger vision. It’s magnetic—it can pull together more brilliant people, creating stronger cohesion. That gives us an edge in organization, and we leverage that edge. It’s a kind of dimensionality reduction attack.

But if you’re just a pure commercial company, and your only vision is to serve C-end users well, you might have an edge in products and traffic—but not in tech. And right now, the most advantageous position is that model technology is the core, the most important thing.

4. We have to be very restrained in many ways. But what’s our core interest? Only one thing: our biggest, even only, core interest is keeping the team stable.

As long as I can keep the team together, we’ll definitely succeed—we’ll definitely achieve AGI. It’s just a matter of sooner or later. If setbacks come and nobody leaves, I can keep going.

5. We’ve always been very restrained. We don’t want to be rivals with any big or small internet company. I hope we can empower everyone, or at least help out. Under that premise, we’re more than willing to assist anyone—even our competitors, like Alibaba, Zhipu, or Moonshot AI—to do better.

Because we don’t lose anything. We’re open-source by nature. If you can’t replicate our work, just ask, and I’ll tell you how. That’s what open source is about. It doesn’t matter if you’re a competitor.

6. Our management runs on two tracks: one top-down, one bottom-up. Bottom-up means everyone decides what they want to work on. They do it on their own, no one micromanages, no KPIs. Top-down is for formal projects that need full company coordination, like launching V4.

We hope formal work doesn’t eat up more than half of anyone’s time. The other half is unscheduled—they can explore whatever they want, as long as the company can support the computing power. No need to ask for permission.

7. We generally don’t work overtime much. Two reasons: First, research needs a relaxed environment. If you push too hard, people can’t do research. Since research relies on your own curiosity and mulling over problems in your own time, you need that relaxed vibe. Second, we’re super focused.

Because we’re restrained, we choose not to do a lot of things. With fewer tasks, everyone has less work. You’ll notice many of our products aren’t polished, and we’re not rushing to fix them. That’s part of our culture.

8. People often say the problem is a lack of talent. But I think that shortage is temporary—it happens at the start of every industry. Remember when we were building websites and internet servers? Talent was scarce then, too. But that shortage gets solved fast, usually within two or three years, because the industry trains up a ton of people.

9. We don’t have anyone to copy. Every step we take is based on real-world analysis of pros and cons, not imitation. I think we’re different from Bell Labs—they clearly didn’t need to commercialize. We do.

At the end of the day, we have to survive. We’re still a company. The government won’t give us a dime. We have to figure out how to stay alive. Many great companies have a pursuit beyond profit, but that pursuit doesn’t hurt their commercialization—it actually makes it better.

Large Model Tech Roadmap and AGI

1. From a tech perspective, the AGI roadmap is actually pretty clear. We can think of AI’s development as a staircase. Last year’s step was CoT (Chain of Thought). We found that with CoT, you can push intelligence to a higher level, raise the ceiling.

This year’s step is the Agent. We’ve realized that with Agents, AI’s scope of capability expands, and the intelligence ceiling goes even higher. Whyis it a staircase? Because each step builds on the one before. Agents use CoT, CoT uses the language model underneath. No step is wasted.

Given the current situation in China, I think the smartest move is to go all-in on general-purpose Agents. Everything else—like finance or medical agents—should be lower priority. Start with coding, because a coding Agent can do a ton.

2. But even after Agents, AI still can’t replace human employees. There’s one thing missing: continuous learning. Humans can learn continuously. You hire a new employee, they spend two months getting familiar with the company and the work, and then they’re up to speed. You say, “Get Xiao Wang,” and they know who Xiao Wang is. But if AI hasn’t had those two months of learning, you’d have to feed it all the context—who Xiao Wang is, their role, where they sit—which just isn’t realistic.

So, after Agents, the next bottleneck we see—and the problem we have to solve—is how to make models learn continuously.

3. After continuous learning, we might hit a singularity. That singularity is: once the model can learn continuously, it can do everything a human can.

It can develop its own versions, do its own research to build the next, stronger AI model—self-iterating. After that step, I think we get embodied intelligence. Then it steps into the real world—doing chores for you, taking care of you in old age. That’s our speculation.

4. We’ve been working on multimodal all along. For C-end user products, it’s important. But for the intelligence ceiling, it’s a component, not the main track. We’ll definitely do multimodal, and future versions will support native multimodal. But we don’t see it as intelligence itself. The main track is like search—search is a component, multimodal is just another component.

5. AI is a broad field. There’s a lot we think isn’t on the main track of intelligence. For example, 3D and video generation—I feel they don’t have much to do with the core intelligence track. We won’t work on them. World models? Same thing—right now, they don’t seem related to raising the intelligence ceiling. So we won’t do that either.

When Sora came out, every big and small company jumped on it. But later, the small ones all cut it. It might be a good business, but it has nothing to do with intelligence. We only do things that are on the intelligence roadmap.

6. We believe in scaling. The bigger the scale, the better the results, the more features you unlock. What’s stopping us from scaling is just compute power. It’s not that we don’t want to scale—we just don’t have enough compute to do it.

We haven’t hit the ceiling yet. When Silicon Valley says scaling has hit a wall, that’s for them. For us in China, we’re nowhere near that point. We haven’t scaled nearly enough. So we’ll spare no effort to push the scaling limit.

7. On data labeling, it comes down to our capital structure. With the way we’re funded, we can’t afford the high-quality data labeling costs they have in the U.S. It’s just too expensive, whether we outsource or do it in-house.

So right now, we’re walking on two legs. We start with the cheaper labeling. But solving AI at this stage is all about labeling data. You could say that half of our core researchers—the most important people—are labeling data. That’s where we’re focused.

8. Internally, we’re really into this narrative: training our next model to help our own development, to boost DeepSeek’s efficiency.

The first goal of the models we build isn’t for you to use well—it’s for us to use well. If we crack continuous learning first, then general intelligence is a piece of cake. With AI’s continuous learning as an aid, it can massively improve our own research efficiency.

Open Source, Commercialization, and Low-Cost Strategy

1. On open source, we were crystal clear from the start. First, it’s the vision. Second, we believe that to make AI work commercially, open source is beneficial.

AI is big enough that it might eventually account for ten percent of human society’s GDP. That’s such a huge number that no one can monopolize it. You have to share, or you won’t survive. It’s different from open-sourcing a regular piece of software, because that market isn’t nearly as big.

If we tried to hog all the benefits, history would leave us behind. You need restraint. You have to build mechanisms to ensure you only take a limited share of the pie, or you won’t succeed.

2. The more restrained you are, the more likely you are to pull this off. Restraint is a strategy. Sometimes, you give up something to get something else bigger. AI is just too vast, and the rewards are too massive. As long as you succeed, the payoff is enormous. Even if you take just a tiny slice, it’s more than enough.

So we’ve said before: we only aim for a reasonable profit. It’s about your intention, not the size of the profit. That’s different from ordinary business thinking.

3. For our API pricing, we think a reasonable profit is: buy a batch of equipment, recover the cost in ten months. That feels right to us.

At that cost, we can break even in ten months. Other companies can’t. Because Alibaba or Tencent lack our optimizations, their costs are probably several times higher. There’s a ton of optimization work involved.

4. In this price range, user demand is inelastic. If you double the price, token consumption barely changes. If it’s twice as expensive, our total revenue nearly doubles. But initially, we worried about too much demand, so we set the model price high. The team wasn’t happy. Then I dropped the price to a quarter of that, and everyone was thrilled.

I think that’s our real mindset. Make a reasonable profit, but make it affordable for everyone. When we cut prices, the company chat group was cheering. Because that’s the point of all the effort we put into building a great model—so everyone can use it fully.

5. Last Spring Festival, we saw a huge surge in users. But we didn’t chase after them to retain them or monetize them. We didn’t fight for commercial gain from users. We just worked hard to serve them well.

We have no intention of building the next super app, or competing to be the next ByteDance or Tencent. I’m not fighting for that. Because there’s a bigger watermelon ahead—what’s in front might just be a few little sesame seeds. Why grab every sesame seed?

6. Last year, it was the C-end. This year, it’s the B-end. I think we should do it and do it well, but that’s not our goal. Most people in our company don’t see this as being on the same level of importance as AGI.

C-end and B-end are both byproducts of our journey toward AGI—intermediate outputs. I’m not doing C-end or B-end for their own sake. I’m doing AGI, and this stuff just happens to come out. So I’ll use it for commercialization.

7. Open source has zero impact on our business model. I’m not worried about others deploying our models to compete with us—not in the slightest. I’m only worried they can’t deploy it properly, miss some details, get worse results, or have higher costs. There’s no conflict here.

8. What were people talking about six months ago when they discussed commercialization? It was always about ads, or e-commerce, or embedding e-commerce into your product and tying it to local services. That’s useless, because things change too fast. If you spend a lot of time on a product’s commercialization path, its lifecycle will be very short.

I don’t think the time is right yet. Over the past three years, any time someone talked to me about a commercialization path or product line, I’d say it was a waste of time. Because you can’t predict the future. Focusing on products too early is always premature. The biggest thing is still the extension of technology.

On Compute, the US-China Gap, and Infrastructure

1. The gap between us and the U.S. is mainly in resources, not people. The people are the same—some are Chinese, some stay in China, some go abroad. Those who go abroad aren’t necessarily smarter; it’s kind of random.

Plus, our base is huge. We have so many new people every year. Talent isn’t the bottleneck; resources are. Resources affect talent development first—because with less compute, we have fewer opportunities to run experiments. So overall, our talent lags behind the U.S. That talent gap is essentially a compute gap.

2. There are too many companies in China building base models. In the U.S., there are maybe three doing it, with resources concentrated. Here, resources are scattered, so each company gets even less. That definitely needs to consolidate—it’s kind of wasteful. We don’t need so many. Right now, people think this is a high-margin business, so they insist on doing it themselves.

But when they realize it’s not that profitable, they’ll stop. I don’t believe in windfall profits—it goes against objective laws. In the end, China will have maybe three or four players competing, and the price will be low enough for a price war. That’s probably enough.

3. In the endgame of large model competition, the differences should come down to three things: cost, time, and user experience. Beyond that, there might not be much else.

Cost is definitely a differentiator. I think cost is the number one factor. Second is time—when you can deliver. Being a few months earlier or later makes a huge difference. Third is experience—there’s some user stickiness there, but it’s probably not fundamental.

4. NVIDIA’s CUDA moat is eroding fast. I think there are three reasons. One is that with AI now, building an ecosystem is much easier than before, because AI can write code.

Second, there are new technologies. For example, we developed something called TileLang—a high-level language. Using it to write CUDA kernels, you can quickly rewrite NVIDIA’s entire ecosystem. Combine that with AI, and there seems to be no obstacle.

Third, the market for compute cards is already bigger than gaming cards. In the future, compute chips won’t be coupled with CUDA anymore—they’ll all be specialized chips. In that context, NVIDIA’s ecosystem advantage shrinks dramatically.

5. There’s a historic opportunity for domestic AI chip replacement. We believe that within a year, we’ll be able to prove that the domestic chip ecosystem is perfectly fine. People used to think it was problematic—that you couldn’t use it, or it wasn’t good. But I think within a year, we’ll change that perception with real results.

The hardware and ecosystem of domestic AI chips are both fine. The only problem is insufficient production capacity. There’s no barrier to adapting domestic cards. NVIDIA can’t stop it. When you can’t buy NVIDIA cards, everyone is forced to use domestic chips. Under that pressure, adaptation has no obstacles—it just takes time.

6. Huawei can supply us with about 16,000 cards. Internet giants might get over 100,000. Those 16,000 Huawei 950 cards are equivalent to about 4,000 NVIDIA B-series cards. So it’s not a huge amount. It’s only enough to train our current generation of models, not the next.

Huawei cards will definitely have a shorter lifecycle, because they’re already two years behind NVIDIA. The Huawei 950 is good to use this year, still okay next year, but if you use it the year after that, it’ll probably be too power-hungry. But I’m optimistic about domestic compute. NVIDIA is digging its own grave.

7. How long until AGI? Can domestic hardware catch up by then? I think in AI, China should be able to match the U.S. within a year or two. Or maybe even this year, we can produce models that are a direct substitute for foreign ones. So that should happen this year—but it’s not AGI yet.

The Huawei 950 super node can fully replace NVIDIA’s GB200 or GB300 in performance and price. The price will definitely be higher, but not by much. Even if it’s 100% more expensive, I think you can consider it a price substitute. Any task the GB300 can do, the Huawei super node can do.

8. Our compute disadvantage is a fact. We deal with it in three ways. First, we accept that our models will lag behind. We have to use smaller models. But that lag has a silver lining: it gives you more time. If you have some technical skill, you can use clever methods.

So our gap with the U.S. is probably 12 to 18 months. Simply put, we’re two years behind, but we do it with one-twentieth of the compute. In the future, we want to rewrite that story: use a fraction of the compute, but shrink the time gap to six months, then three months.

9. On the process and decision-making mechanism for major strategic and technical decisions: The company is built on consensus. I don’t decide everything alone. I seek consensus. My authority and influence within the company are based on that consensus.

For example, if I want to do something, I first see what the consensus is—what everyone wants to do. Then I might guide or lean a certain way, but that guidance is very limited. It has to be built on consensus for me to push it through.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

Log In / Sign Up

Enter your email to receive a secure code. No password needed.