I’ve been spending a lot of time using local LLMs lately, and most of the time the process is boredom-inducing. Load the model, run a CLI command, nod at the tokens-per-second figure as if I understand what’s going on. Those numbers only tell part of the story, as knowing how fast something is doesn’t tell you how long it can stay on task.