Which Local LLM is the Best for the RTX3060? (26 Candidates, 1 Winner)

The video evaluates 26 local large language models on an Nvidia RTX 3060, finding Google’s Gemma 3 12B to be the best overall balance of speed, quality, and VRAM efficiency for this hardware. While other models like Ling Mini and IBM’s Granite excelled in speed and tool use respectively, and Qwen 3 30B showed top math reasoning performance with CPU offloading, Gemma 3 12B is recommended as the most versatile choice for typical RTX 3060 users.

The video presents a comprehensive showdown of 26 local large language models (LLMs) tested on an Nvidia RTX 3060 GPU with 12GB VRAM, aiming to identify the best-performing model for this specific hardware setup. The host outlines the testing environment, emphasizing the constraints of VRAM and the use of llama.cpp with CUDA for efficient model execution. The models tested include a mix of popular open models, lesser-known dark horses, and some wild cards, all evaluated on speed, quality, VRAM fit, and the ability to use speculative decoding to enhance performance without sacrificing output quality.

The evaluation process involved running each model through a series of benchmarks covering math reasoning, coding ability, instruction following, and agentic tool use, with 100 samples per task. The host developed a custom evaluation harness, LM Eval, to automate and standardize testing across all models. The testing took several days due to the extensive number of models and the computational demands, with some challenges encountered in scoring, especially for models with complex tool use capabilities or those requiring CPU offloading.

Results showed that in math reasoning, the Qwen 3 30B A3B model, a mixture of experts (MoE) model, performed best but required CPU offloading, making it less practical for everyday use on the RTX 3060. For coding tasks, the Qwen 2.5 coder and Google’s Gemma 3 12B tied for the top spot, with the Ling Mini model also performing impressively despite its sparse activation. In agentic tool use, IBM’s Granite 4.18B was a surprising leader, excelling in tool calling capabilities. Speed tests crowned the Ling Mini as the fastest model, nearly doubling the speed of others while maintaining strong performance.

The overall winner of the showdown was Google’s Gemma 3 12B model, which balanced speed, quality, and VRAM efficiency, fitting comfortably within the 9.5GB VRAM limit of the RTX 3060. While not the absolute best in any single category, Gemma 3 proved to be the best all-around performer across all tested axes, including reasoning, coding, instruction following, and tool use. The host notes that although Gemma 4 is a newer and more advanced model with enhanced reasoning capabilities, its benefits did not fully manifest in the specific benchmarks used, partly due to the need to limit its thinking mode for evaluation.

In conclusion, the video recommends Gemma 3 12B as the top local LLM for users with an RTX 3060 and around 16GB of system RAM, highlighting its versatility and efficiency. The Ling Mini and Granite 4 models are also recommended for users prioritizing speed or tool use, respectively. For those willing to trade speed for quality and have more powerful hardware, the Qwen 3 30B MoE model is suggested. The host invites viewers to share their experiences and suggests potential future tests including more customized or fine-tuned models, encouraging community engagement and ongoing exploration of local LLM performance.