THE HONEST CALCULATOR
CAN I RUN IT?
Pick your machine. We check it against the channel's model-research database — real quant file sizes, real measured speeds — and tell you what actually fits, what fits only by spilling into system RAM, and what the hype conveniently left out. Where there's no measured number for your rig, we say so instead of guessing.
Unified-memory machine (Mac, Strix Halo, GB10)? Put your total memory in VRAM and leave RAM at 0.
WHAT FITS
What fits your rig
NOT ON YOUR HARDWARE
Models the hype says you can run at home. You can't — and here's the receipt, the same on every rig.
HOW WE KNOW
Every number on this page is pulled straight from the channel's own model-research database — the same notes behind the videos — last verified the database date. Nothing here is invented.
- Fit is a disclosed rule, not a vibe: a model "fits with headroom" only when its file takes at most 90% of the memory you selected AND the research notes don't name a higher minimum — the notes always outrank our arithmetic. A file that squeezes in without meeting both tests shows as "borderline", and the research note on the card is the real verdict. "With CPU offload" appears only when the research says spilling into system RAM actually works.
- Real use needs headroom. The weights are only part of it — context and the KV cache also live in memory and grow with how much you feed the model. Treat a tight fit as "yes, with a modest context", not "yes, with the full 256K".
- Speeds are measured, not modelled. We only show tokens-per-second when the research has a real result on similar hardware, with its source rig and quant attached. Extrapolations and vendor claims are left off on purpose.
- This is guidance, not a guarantee. Your OS, drivers, runtime (llama.cpp, MLX, vLLM…), context length and quant all move the real number. Use it to shortlist, then try it.
$ source → channel model-research DB · db verified — · page generated —