Muse Glimmer 30B runs on my 32GB GPU, handles its own tool calls, and never leaves the machine
I've spent most of this year trying to get the local models I run in a Proxmox LXC to behave like an agent rather than a chatbot with a to-do list, and the results have been mixed. Small models call the wrong tool; medium models call the right tool and then panic when it errors out; and the big models that can cope don't fit on a single consumer GPU. So when Meta dropped Muse Glimmer 30B under Apache 2.0 with the pitch of "always-on local agents," I was interested and skeptical in equal measure.
I’ve spent most of this year trying to get the local models I run in a Proxmox LXC to behave like an agent rather than a chatbot with a to-do list, and the results have been mixed. Small models call the wrong tool; medium models call the right tool and then panic when it errors out; and the big models that can cope don’t fit on a single consumer GPU. So when Meta dropped Muse Glimmer 30B under Apache 2.0 with the pitch of “always-on local agents,” I was interested and skeptical in equal measure.
Clarence Harper
United States
United States
Published by: aplhsindia.in
