I tested Gemma 4, Qwen 3.5, and Ministral 3 for vision tasks, and only one understood the a**ignment
Local models get put through their paces on plenty of fronts. Reasoning, general chat, mobile-friendly setups, whatever fits a small card; all of that gets tested to death (including by me). Vision is one of the areas that doesn't come up as much in these head-to-heads, which is odd because for a lot of daily local LLM use it's the capability that matters the most. I use vision constantly for screenshots I need help with or just random photos of things I can't identify.
Local models get put through their paces on plenty of fronts. Reasoning, general chat, mobile-friendly setups, whatever fits a small card; all of that gets tested to death (including by me). Vision is one of the areas that doesn’t come up as much in these head-to-heads, which is odd because for a lot of daily local LLM use it’s the capability that matters the most. I use vision constantly for screenshots I need help with or just random photos of things I can’t identify.
Michael Johnson
Chicago
Chicago
Published by: aplhsindia.in
