When a local model picks its next word, it doesn't simply know the answer. It scores every possible word at once, handing you a whole spread of candidates, each glowing a little brighter or dimmer depending on how likely it is.

Sampling is just how it reaches into that spread and chooses one. If it always grabbed the single brightest candidate, your text would turn stiff and repetitive, so instead it picks with a little controlled chance.

But the full list of candidates is enormous, so we trim it first. Top k keeps only the few most likely words, while top p keeps just enough of them to cover most of the probability, then throws the long tail away.

Temperature is the heat, not the choosing. Turn it up and the candidates spread wider and wilder, turn it down and they tighten toward the safest few, and then sampling reaches in and grabs one.

So the shortlist narrows, the heat sets the mood, and one candidate flares as the chosen word. Repeat that a few dozen times and you get a sentence, one lucky pick at a time. That's sampling. Off the Cloud.