Ever since I got into hosting my own large language models, I’ve tinkered with AI inference engines across multiple devices. So far, my RTX 3080 Ti remains my primary choice for deploying the Qwen3.6-35B-A3B for coding-heavy tasks using MoE offloading, though I’ve also had decent experience using Gemma-4-26B-A4B and GPT-OSS-20B on old Pascal-era graphics cards.