News Reporter |
News Editor | Wen Shuqi
No model maker can afford to let their guard down these days. Within a span of roughly 48 hours, DeepSeek and Zhipu both dropped their latest model updates back-to-back.
First up, DeepSeek updated its API documentation to the DeepSeek-V4-Pro official version on the evening of August 12, then pulled back and re-released the announcement on the 13th—though API services stayed live throughout the adjustment window. Then, on August 14, Zhipu officially unleashed GLM-5.3.
Leaving aside the messy rollout drama around DeepSeek’s retraction (plenty of third-party testers at the time claimed results didn’t match the official hype, some even saying the Flash version felt better), the official word is that DeepSeek V4 Pro is built with Agent capabilities front and center, showing noticeably sharper performance gains in real-world production tasks.
On Agent-focused benchmark suites, its overall scores land right alongside Kimi K3, Opus 4.8, and Fable 5.
Zhipu’s GLM-5.3, meanwhile, doubles down on coding chops. This model shares the exact same base model as its predecessor GLM-5.2, but gets its boost purely from post-training upgrades—using reinforcement learning built on IndexShare, SAO, and the new Slime framework to stretch long-horizon task environments and training time.

Looking at benchmark performance, GLM-5.3 edges close to Fable 5 and GPT-5.6 Sol in certain areas, and actually beats Kimi K3 across several published leaderboards—though it still trails K3 on DeepSWE v1.1, which zeroes in on long-horizon software engineering and sustained code modification.
Security is the other headline feature for GLM-5.3, a topic that got some buzz earlier when Anthropic’s Mythos 5 stirred the pot on model safety.
According to the company, in CyberGym testing—where the model starts from white-box source code and hunts for vulnerabilities by triggering program faults—GLM-5.3 scored 84.5%, nosing past Mythos 5’s 83.8% and GPT-5.6 Sol’s 83.6%. But on ExploitBench, which demands deeper reasoning, GLM-5.3 pulled 54.4%, well below Mythos 5’s 78.0% and GPT-5.6 Sol’s 76.5%.
On the user side, possibly because of the early hiccups with DeepSeek-V4-Pro, several testers who ran both models told us that GLM-5.3 simply felt better than DeepSeek-V4-Pro when it came to coding tasks.
That said, one developer made it crystal clear he wouldn’t make either his daily driver. “Among domestic models, K3 is still the only one in the top tier.” He added that it’s not just about well-rounded capability gains—the cost-effectiveness matters just as much.
That verdict lands against a big backdrop: DeepSeek just overhauled its API pricing.
The API is introducing peak-valley pricing for the first time, with off-peak usage priced at half of peak rates. But here’s the kicker—comparing DeepSeek V4 Pro’s peak pricing to before, cache-hit input costs jumped 1,100% year-over-year, cache-miss input went up 200%, and output climbed 350%.
Zooming in on the K3 versus DeepSeek-V4-Pro gap: before the hike, K3’s cache-hit input price was roughly a thousand times DeepSeek’s; now it’s around ten times. Cache-miss input went from over 6x down to about 2x at peak hours, and output pricing dropped from roughly 16x to about 4x.
The once-lopsided value equation has shifted. Now that you factor in conversation turns, output quality, and total task duration, developers with different needs might arrive at very different answers when choosing.
Here’s the thing—both DeepSeek-V4-Pro and GLM-5.3 are essentially post-training iterations on top of last-gen base models. Kimi K3’s breakthrough, by contrast, came from scaling up base model training even further. So what do these two different paths actually tell us?
One AI-focused investor told us that, from a technical standpoint, it proves two things: pre-training scale-up still matters, and post-training at existing scale has plenty of headroom left. The better proof point here? Grok 4.6, a 1.5T-parameter model, nearly matches GPT-5.6 and Opus 5.
For domestic model makers, this means true generational leaps probably still require scaling up the base model. And until they get there, Kimi K3’s potential clearly hasn’t been fully tapped yet.