Every time you ask a local model something and it answers, that's inference. Training was the long, expensive part that built the brain, and inference is just switching it on and putting it to work.
Picture the finished model as a settled glowing sphere. Training spent all that time and energy shaping it, but once it's done the sphere simply sits there, steady and ready to go.
Ask a question and the model wakes up for a moment. It reads your words, then flows out an answer one piece at a time, each new word guessed from everything that came before it.
Here's the nice part. Inference is light enough to run right on your own machine, on your GPU or even your CPU, with no cloud, no bill, and no one watching.
The expensive part, training, only happened once, long before you ever downloaded the model, so from here you can run it as often as you like on your own hardware. That's inference. Off the Cloud.