DEEP DIVE · EPISODE
DeepSeek: The Whale That Ate the Stock Market — Can You Run It?
The weights are free. The question this channel exists to answer is whether you can actually run them. A tour of the whole DeepSeek dynasty — from the R1 earthquake to the V4 monster — measured, priced, and weighed in VRAM.
$ play --episode → 4:57 · open weights, weighed in VRAM
The whale that arrived
In January 2025 a Chinese lab that nobody's parents had heard of released a reasoning model with open weights — and within a week NVIDIA had suffered the largest single-day value wipeout in stock-market history. The whale had arrived. Since then the DeepSeek family has shipped model after model, every one open weights, every one making headlines, and every one raising the same question: the weights are free, but can you actually run them? This episode weighs the whole dynasty — from R1 to V4 — measured, priced, and counted in gigabytes of VRAM.
For the record, the R1 moment was a frontier-class reasoning model — the thinking-out-loud kind — dropped under an MIT license, with a paper explaining how it was built. The panic wasn't just that it was good. It was that it was good, open, and allegedly cheap to train. And then came the legend that launched a thousand thumbnails: R1 runs on your gaming PC, thanks to a magical 1.58-bit quant. A 671-billion-parameter model, on your graphics card. You can already guess how that ends.
The 1.58-bit legend, weighed in a real room
Here is what that legend actually looks like once you put it in a room with real hardware. The famous dynamic 1.58-bit build is 131 gigabytes, so a 24 GB card holds less than a fifth of it and the rest sits in system RAM. Pair a 24 GB GPU with 128 GB of RAM and you get maybe two to five tokens per second. Now remember what kind of model this is: a reasoning model that thinks in thousands of tokens before it answers. At five tokens a second, a hard question becomes a lunch break. And the quant itself paid for the squeeze — on the quantizer's own test suite it scored 69 percent, against roughly 92 for the larger builds. It technically runs. In practice it's a space heater with opinions.
The real gift: the distills
But the whale left a genuine gift on the beach: the distills. DeepSeek compressed R1's reasoning into normal-sized models, and the 32B distill at Q4 fits a single used RTX 3090, where it does a real 20 to 25 tokens per second while keeping a genuine chunk of the reasoning ability. That is several times faster than torturing the full whale onto the same card, at a quality that is actually usable. This was the real consumer story of R1 all along: not running the giant, but running its student. The giant was always a datacenter animal wearing a free sticker.
The efficiency era: V3.2
Act three is V3.2 — the efficiency era. Sparse attention, 671 billion parameters, genuinely brilliant engineering, all of it aimed at datacenters. The smallest quant anyone has built comes in at 161 gigabytes. That fits no consumer setup — not a single 5090, not two of them. Even testers stacking NVIDIA's 96 GB workstation cards, in rigs that cost as much as a house deposit, report speeds in the low double digits. The whale had simply grown. API-efficient is not the same as home-efficient, and 37 billion active parameters still need all 671 billion sitting in memory. Say it with the channel: active parameters save you time, never space.
V4: open weights, closing doors
Then in April came V4, which earned a whole episode of its own. The short version: Flash — the "small" one — is 284 billion parameters with an 81-gigabyte entry ticket on exotic forks, while Pro is 1.6 trillion and basically a mortgage. By now the family pattern is unmistakable. Every generation the models get smarter, the licenses stay open, and the entry ticket gets bigger. R1 asked for around 130 gigabytes. V3.2 asked for 160. V4 Flash asks for 81 on exotic forks, and Pro asks for the better part of a terabyte. Open weights, closing doors.
The DeepSeek verdict
So here is the verdict. What you should actually run is the R1 distills — the 32B on a 24 GB card is still one of the best local reasoning deals of 2026. What you should admire from the shore is everything else. The whales live in datacenters, and at DeepSeek's famously cheap API prices, renting the whale costs less than the electricity you'd burn torturing it locally.
The bigger lesson the whale taught local AI is that open weights and runnable weights are not the same thing. DeepSeek proved a lab can give away frontier intelligence and still shake the world — and it also proved that a Hugging Face link is not a hardware plan. The gap between "free" and "yours" is measured in gigabytes, and it keeps growing. That's the DeepSeek dynasty: one earthquake, one gift, and a family of whales too big for any tank you own. Stay local, stay skeptical.
THE RECEIPTS
Every figure below is stated in the episode — open weights, weighed in VRAM.
- DeepSeek R1 — 671-billion-parameter reasoning model, MIT license, open weights (January 2025).
- R1 dynamic 1.58-bit build — 131 GB. A 24 GB card holds under a fifth of it; the rest sits in system RAM.
- R1 1.58-bit on a 24 GB GPU + 128 GB RAM — roughly 2 to 5 tokens per second.
- That 1.58-bit quant's quality — 69% on the quantizer's own test suite, versus about 92% for the larger builds.
- R1 32B distill at Q4 — fits one used RTX 3090 (24 GB), around 20 to 25 tokens per second.
- DeepSeek V3.2 — 671B parameters; smallest quant built is 161 GB; fits no consumer rig.
- DeepSeek V4 Flash / Pro — Flash is 284B parameters with an ~81 GB entry ticket on exotic forks; Pro is 1.6 trillion parameters.
- The entry ticket, over time — R1 ~130 GB → V3.2 161 GB → V4 Flash 81 GB → V4 Pro the better part of a terabyte.
WATCH & SUBSCRIBE
Watch the full episode on YouTube — and a new local-AI hype-check drops on the channel every day, receipts included.
Wondering whether the R1 distill fits your card? Check your rig →