DEEP DIVE · EPISODE
GLM 5.2 & Kimi: Free Giants You Can't Run
Two Chinese labs took the open-weights crown in 2026. This is why you still can't run either one — and the two models on the same shelf that you actually can.
Runtime 5:35 · click the frame to play
The best open model on Earth right now is GLM 5.2. It is MIT licensed, completely free, the number one open-weights model on the biggest intelligence index there is, and it trades blows with DeepSeek and Kimi on the coding boards. And unless you own a quarter-terabyte of memory, you will never run it. This episode is about the two Chinese giants — GLM and Kimi — that took the open-weights crown in 2026, and the uncomfortable truth underneath: there are now two leaderboards in local AI. One measures how good models are. The other measures what you can actually run. The gap between them is the story of the year.
GLM's résumé is real
GLM 5 shipped in February 2026: 744 billion parameters, MIT license, trained on 28.5 trillion tokens. GLM 5.2 followed in June and took the number one spot on the Agent Arena leaderboard, posted the highest open-weights intelligence score ever recorded, and landed within about one percent of Claude on a major software-engineering benchmark. One security firm even reported it beating Claude on their cyber tests. This is not benchmark cosplay. The open-weights frontier now speaks Mandarin, and it is genuinely frontier.
GLM's hardware bill
Then comes the invoice. The recommended two-bit build weighs 239 gigabytes and scores about 82 percent on Unsloth's own token-accuracy test — and the smallest quant anyone has made, at one bit, is still 217 gigabytes. That fits no consumer card. Not a 5090, not dual 3090s, not a 128-gig Mac. The entry ticket is a 256-gig Mac Studio grinding out three to nine tokens per second — a whisper from a mountain. But GLM left a real gift: GLM 4.7 Flash, 30 billion total parameters with 3 billion active, MIT licensed, about 18 gigs at Q4, running 60 to 100 tokens per second on a used 3090. The crumb is actually a meal.
Kimi builds whales
Challenger two is Kimi, from Moonshot AI, and the whole family line is a flex. K2, K2.5, K2.6, and the coding-focused K2.7 Code — every single one is a trillion parameters, 32 billion active, open weights under a modified MIT license. The coding claims are serious: 80 percent on the big software benchmark. The rumor mill says K3 will be two and a half trillion. Kimi does not make small models for your amusement. Kimi makes whales and dares you to find an ocean.
Kimi's hardware bill
Same story, bigger numbers. The smallest quant that isn't broken is about 247 gigabytes; the recommended two-bit runs around 385. The famous "it runs on a Mac" line? That is the $9,500, 512-gig Mac Studio — a configuration Apple quietly pulled this spring — doing five to twenty-odd tokens per second. One well-known tester lashed four Mac Studios together, forty grand of aluminium, to reach twenty-eight. The vendor's "over 40 tokens per second" figure is a datacenter B200 number. For the record, there is one runnable Kimi: Kimi Linear, 48 billion parameters with 3 billion active — 30 gigs at Q4, dual-3090 territory, or three-bit on a single 24-gig card.
The stat of the year
Take the top ten open models by measured quality in July 2026. The only ones that squeeze onto a 24-gig consumer card are the 27-to-31-billion dense pair: Gemma 4 31B, and Qwen 3.6 27B, which actually outranks it. Add a dual-3090 rig at 48 gigs and you climb a little higher — but the true heavyweights, the hundred-billion-plus mixtures of experts like GLM and DeepSeek, still sit well out of reach. The other eight live in server racks and half-terabyte Macs. The quality leaderboard and the runnable leaderboard have almost completely split, and the runnable one is topped by the 27-to-31B class: Qwen, Gemma, and friends. The frontier is open. The frontier is also, for most of us, a museum with free admission and a locked door.
So what actually runs?
The honest verdict, for your hardware: run GLM 4.7 Flash, Kimi Linear, or the Qwen 27B class. That is where open weights and your VRAM actually meet. As for the giants, renting them is cheap but not quite pennies — the cheapest whales serve for under thirty cents a million tokens, while Kimi's output runs a few dollars. Their existence still drags every API price down, and because the weights are public you still get distills, forks, and pressure even if you never load a single shard. "Open to rent" is not the same as "open to run" — but it still beats closed. Just barely. And it still isn't what the thumbnails promised.
Open for whom?
Two leaderboards, one wallet. When a lab says open, the question this channel keeps asking is: open for whom? The answer in 2026 — open for datacenters, open for landlords with half-terabyte Macs, and open, in spirit and in distills, for the rest of us. That is GLM and Kimi: magnificent, free, and further from your graphics card than the moon. Stay local, stay skeptical.
THE RECEIPTS
- GLM 5 (Feb 2026): 744B parameters, MIT license, trained on 28.5 trillion tokens.
- GLM 5.2 (Jun 2026): #1 on the Agent Arena leaderboard; highest open-weights intelligence score recorded; within ~1% of Claude on a major software-engineering benchmark; one security firm reported it beating Claude on their cyber tests.
- GLM 5.2 quant: recommended two-bit build is 239GB, ~82% on Unsloth's token-accuracy test; entry hardware is a 256GB Mac Studio at 3–9 tokens/sec.
- GLM 4.7 Flash: 30B total / 3B active, MIT, ~18GB at Q4, 60–100 tokens/sec on a used RTX 3090.
- Kimi (Moonshot AI): K2, K2.5, K2.6, K2.7 Code — each 1 trillion parameters, 32B active, open weights, modified MIT; 80% on the big software benchmark; K3 rumored at 2.5 trillion.
- Kimi quant: smallest usable ~247GB; recommended two-bit ~385GB; the "runs on a Mac" claim is the $9,500 / 512GB Mac Studio (pulled spring 2026) at 5–20+ tokens/sec; four lashed-together Mac Studios (~$40k) reached 28; the vendor's 40+ tokens/sec figure is a datacenter B200 number.
- Kimi Linear: 48B / 3B active, ~30GB at Q4 (dual 3090), or three-bit on a single 24GB card.
- Top ten open models (Jul 2026): only the 27–31B dense pair fit a 24GB card — Gemma 4 31B and Qwen 3.6 27B, which outranks it; the hundred-billion-plus MoEs like GLM and DeepSeek stay out of reach.
- Rent economics: the cheapest giants serve for under $0.30 per million tokens; Kimi's output runs a few dollars per million.
Figures as stated in the episode.
WATCH · SUBSCRIBE
Want to know which of these fits your box? Run GLM 4.7 Flash, Kimi Linear and the Qwen 27B class through the Can I run it? calculator.