Despite my initial scepticism, locally-hosted LLM models have gotten a lot of oomph to their reasoning capabilities as of late. Newer Mixture-of-Experts models, for example, can run at respectable token rates on my outdated Pascal-era cards, and with a little bit of tinkering, LLMs such as Qwen3.6-35B-A3B and Gemma-4-26B-A4B can easily replace their cloud counterparts for my FOSS stack.