Ever since I used my aged GTX 1080 to build a local AI hosting workstation last year, I’ve benchmarked LLM inference tasks across all sorts of hardware, ranging from gaming systems to tiny single-board computers and outdated mobile phones. As you’d expect, devices with NPUs and VRAM-laden graphics cards work best with my local LLMs, and with the right tweaks, can even drive bulky MoE models.