Over the last couple of months, I’ve tried running local LLMs on several servers, ranging from full-on RTX cards, NAS units, and old gaming PCs to random laptops and Raspberry Pi boards. To no one’s surprise, I’ve had the best luck with Nvidia cards, and not just modern ones featuring Tensor cores. In fact, I currently use an old GTX 1080 to run Gemma-4-E4B for my everyday productivity Docker apps, and this behemoth of a card can even drive the likes of GPT-OSS-20B and Gemma-4-26B-A4B at respectable speeds as long as I rely on MoE offloading.