DEEP DIVE / EPISODE

What Can YOUR GPU Actually Run? The July 2026 Buyer's Guide

GPU prices went mad, DRAM is up around 98% in a quarter, and the internet still says just buy a 5090. Here's what every card actually runs in July 2026 — measured, priced, no wishful thinking.

What Can YOUR GPU Actually Run? The July 2026 Buyer's Guide

$ episode --runtime → 5:08 · long-form buyer's guide · measured on named hardware

First, the crime scene

Memory prices went vertical this year. DRAM contract prices climbed as much as 98% in a single quarter, and everything downstream repriced along with them. The RTX 5090 carries a $1,999 MSRP but streets around $3,000, with premium cards pushing well past $4,000. Used 3090s — the people's card — stubbornly hold around $1,000, and the trackers show them creeping up while everything else goes vertical. Apple raised Mac prices across the board and quietly killed the 512 GB Mac Studio, the one machine that could run the biggest local models. NVIDIA's little DGX Spark box was hiked to about $4,700. The rule that follows from all this is simple: if a buyer's guide is still quoting 2025 prices, close the tab.

8 GB — the laptop tier

Down here there's really one answer: Qwen 3.5 9B at Q4. It fits a full 32,000-token context inside an 8 GB card, which nothing else its size manages. The iron rule at this tier is that the model must live entirely on the GPU. Spill even a handful of layers to the CPU and on many cards your speed can collapse by 70% or more — the difference between an assistant and a slideshow. If it doesn't fit, don't force it. Go smaller and stay fast.

12 GB — the people's tier

This is home turf: the very episode was rendered on two 12 GB cards. The used RTX 3060 still sells for about $230, and it runs a 14B at Q4 at roughly 23 tokens per second with 16,000 tokens of context — or an 8B at about 32 tokens per second with double the context. That is real, daily-driver local AI for the price of a night out. Pound for pound it is the best value in the entire market, and it is not close.

16 GB — the trap tier

It sounds like more, but it buys you less. The hot new 27B models are 17 GB files at Q4, so they do not fit. Your options are to drop to 3-bit, with the quality already sliding, or to run the older Mistral Small 3.1 — about 14 GB at Q4, 20 to 50-plus tokens per second depending on the card, and still genuinely excellent. And one landmine to watch: the new "Mistral Small 4" is not small. It's a 119-billion-parameter mixture of experts that needs on the order of 70 GB even at four-bit — the name is pure marketing. 16 GB cards are the awkward middle child of 2026.

24 GB — the kingdom

This is the tier of 2026. A used 4090 or 3090 runs almost the entire 27-to-35B class at Q4 — everything but the one dense 31B that maxes the card out. Qwen's 27B lands around 47 tokens per second on a 4090; Gemma's 26B mixture of experts was measured at about 149 tokens per second on the same card, an absolute rocket. And a quiet revolution nobody is shouting about: the new models' hybrid attention cut long-context memory to roughly a quarter of the old cost, so 128,000 tokens of context on a single card just went from fantasy to Tuesday.

32 GB and beyond

The RTX 5090 at around $3,000 buys you 5- and 6-bit quality on the 27B class, but a 70B still only squeezes on at 3-bit, barely. The honest power move remains two used 3090s: 48 GB for around $2,000, running a real Q4 70B at 17 to 21 tokens per second. Two fabrications to dodge while you shop: there is no "Llama 3.3 14B" — that model simply does not exist, no matter what the listicle says — and the "17,000 tokens per second" RTX 5090 claim is throughput-benchmark marketing, off by roughly 100x from the chat speed you will actually see.

The one rule that survives the chaos

The cheat sheet in a single breath: 8 GB, a 9B; 12 GB, a 14B and the best value in the game; 16 GB, the awkward tier, so Mistral or 3-bit; 24 GB, the kingdom and the 27B class at Q4; 32 GB, a quality bump rather than a new world; 48 GB from two used cards, a real 70B. And the golden rule when buying: VRAM first, compute second. A slower card that fits your model beats a faster card that doesn't, every single time. Prices are moving monthly, so check before you click. This is Off the Cloud — what actually runs, at what speed, on hardware with a name. Stay local, stay skeptical.

THE RECEIPTS

Every number in the episode, checkable. Hardware prices are volatile — the figures are dated where the episode dates them.

WATCH IT / SUBSCRIBE

The full buyer's guide, tier by tier, with the footage and the numbers on screen. Five minutes, no filler.

Want this narrowed to your exact card? The Can I run it? tool does the fitting for you.

HYPE, CHECKED — WEEKLY

One email a week: what actually shipped in local AI, what was hype, receipts included. No spam, unsubscribe any time.