Recently, China’s AI powerhouse MiniMax (Shanghai Xiuyu Technology) dropped its latest flagship model, M3, and at the same time pulled the plug on its old Coding Plan—which billed you per request—swapping it out for a new Token Plan that charges based on actual token usage. They also quietly scrapped their cheaper tier. Naturally, this caught a lot of people off guard. Users took to social media en masse to vent, pointing out that they weren’t given enough heads-up, weren’t consulted, and found themselves burning through monthly credits way faster than expected just running the same tasks many discovered the billing rules had quietly changed right when they logged in on June 1st. For a lot of folks, it felt less like a necessary update and more like a sneaky price hike that shrunk the value of what they were already paying for.

With the complaint volume spiking, MiniMax quickly jumped online to issue an apology and lay out what’s changing. In their statement, they admitted they didn’t do a great job keeping everyone in the loop ahead of time, especially when it came to explaining how the new Token Plan would work alongside the M3 model. They acknowledged that handling week limits for long-time subscribers was rushed and could’ve been managed much better.
So why the switch? MiniMax broke it down by pointing out that the new M3 model is genuinely beefier—it’s larger, packs stronger reasoning capabilities, natively handles multiple data types, and boasts a massive 1-million-token context window. Those features let it tackle way more complex jobs, but they also drain compute resources fast, which means the old pay-per-call model just doesn’t cut it anymore. On top of that, the company heard loud and clear from users who wanted the flexibility to spend their subscription credits across different model types instead of being locked into one format. Moving to a standard token-based metric was basically the industry norm, so they decided to align with it.
To make things right, MiniMax rolled out a three-part fix. First, if you grabbed a plan before March 22nd and were enjoying unlimited weekly quotas, that perk stays intact—you’ll still get full access to both M2.7 and M3 without those weekly caps. Second, anyone who signed up for the Token Plan between March 22nd and 10 AM this past Friday gets a permanent 50% boost on their M3 weekly quota for the life of their subscription. Third, to give everyone a proper taste of how M3 handles heavy lifting, the company is resetting all credit balances right now. Plus, between June 1st and June 7th, they’re doubling the standard “5 hours per week” allowance for all subscribers. They’re also stretching out bonus credit expiration dates from one month to a whole year and are pushing hard to launch a self-service refund portal ASAP.
Back on June 1st, MiniMax unveiled the M3, built entirely on their own custom MSA (MiniMax Sparse Attention) architecture. This isn’t just another text generator; it’s a fully native multimodal model that can digest images and video side-by-side while actually navigating your desktop like a human. It’s currently the only domestic model that successfully bundles cutting-edge coding skills, a 1-million-token context window, and true multimodal understanding into one package.
Tied directly to the M3 launch is the new billing setup. The fresh “Token Plan” swaps out the old pay-per-call system for a metered approach where you’re charged strictly by how many tokens you burn. Before this shift, MiniMax’s Coding Plan—built for indie devs and casual users—ran on a simple per-API-call basis. You’d drop a flat monthly fee (say, 49 RMB for the Plus tier), get a set number of calls within a rolling 5-hour window, and never worry about hitting a total token cap. Honestly, that model was incredibly generous for folks tackling marathon coding sessions, deep-dive research, or immersive translation projects.

But honestly, the new Token Plan flips the entire script. Take that same 49 RMB Plus tier, for instance. Instead of an open-ended well, you’re now handed a strict 600 million tokens once a month. The moment that bucket runs dry, you’re stuck until the next cycle. We went from “capped speed, unlimited volume” straight to “strict volume cap, uncapped speed.” Heavy users immediately saw their costs skyrocket. One developer did the math and pointed out that back in the day, burning through 3 to 5 billion tokens a month cost a flat 49 RMB. Under the new rules, getting that same amount of juice now runs you closer to 175 RMB—a jump of over 257%. That’s a tough pill to swallow.
And frankly, MiniMax isn’t alone in feeling the heat. Just recently, another major player, Moonshot AI (the team behind Kimi), faced a similar wave of backlash after tweaking their own billing structure. As these models get sharper—especially when you plug them into AI Agent workflows—a single task can easily trigger dozens, sometimes hundreds, of automatic calls, thoughts, and actions. Token consumption goes exponential overnight. From a business standpoint, sticking to “pay-per-call, unlimited” simply doesn’t fly in the agent era. Compute bills spiral out of control pretty fast. Shifting to granular, token-level pricing is really just the industry catching up to reality—a necessary move to keep costs manageable and stay in business long-term.