I’ve spent most of this year trying to get the local models I run in a Proxmox LXC to behave like an agent rather than a chatbot with a to-do list, and the results have been mixed. Small models call the wrong tool; medium models call the right tool and then panic when it errors out; and the big models that can cope don’t fit on a single consumer GPU. So when Meta dropped Muse Glimmer 30B under Apache 2.0 with the pitch of “always-on local agents,” I was interested and skeptical in equal measure.