Gemini 3.1 Pro is the smartest model ever made

The video reviews Google’s Gemini 3.1 Pro, praising its record-breaking intelligence and benchmark performance but criticizing its poor usability, reliability, and tool integration in real-world development tasks. The creator concludes that despite its advanced knowledge, Gemini 3.1 Pro is frustrating for practical use, recommending more consistent models like Claude for coding and automation until Google improves its user experience.

The video reviews Google’s newly released Gemini 3.1 Pro, which is being touted as the smartest AI model ever made. The creator highlights that Gemini 3.1 Pro has achieved record-breaking scores on multiple benchmarks, outperforming previous leaders like Opus 4.6 Max, and doing so at less than half the cost. Notably, it scored an impressive 78% on the ARC AGI 2 benchmark, a feat previously thought unattainable. The model also excels in niche and complex tasks, such as spatial reasoning and SVG animation, and demonstrates a significant reduction in hallucination rates compared to earlier versions.

Despite its intelligence, the creator expresses deep frustration with Gemini 3.1 Pro’s usability, particularly when integrated into development workflows. The official Gemini CLI is described as buggy, unreliable, and often fails to use the correct model or execute tool calls properly. Even when using alternative interfaces like Cursor, the experience remains inconsistent, with the model frequently failing at basic tasks, getting stuck in loops, or producing nonsensical outputs. These issues make the model difficult to use for real-world coding and automation tasks, despite its theoretical capabilities.

The video contrasts Gemini 3.1 Pro’s intelligence with its lack of practical competence, especially in tool usage and long agentic runs. While other models like Anthropic’s Claude 4.5 Haiku may be less intelligent, they are much more reliable and consistent in following instructions and using tools as intended. The creator argues that Google’s focus on maximizing benchmark scores (“benchmaxing”) has come at the expense of real-world usability, leaving Gemini 3.1 Pro feeling like an outdated model with a massive knowledge upgrade but lacking the behavioral improvements seen in competitors.

The creator also discusses various benchmarks and real-world applications, noting that Gemini 3.1 Pro excels in knowledge-based tasks and can adapt well when given explicit guidelines, as seen in the Convex LLM leaderboard. However, it performs poorly on tasks requiring nuanced behavior or ethical judgment, such as the SnitchBench, where it is overly eager to “snitch” in scenarios involving sensitive information. The model’s tendency to get stuck, hallucinate, or misuse tools means that users must closely monitor its actions, negating some of the efficiency gains from its intelligence.

In conclusion, while Gemini 3.1 Pro represents a significant leap in AI intelligence and knowledge, its practical shortcomings make it frustrating to use for developers and power users. The creator urges Google to shift focus from chasing benchmark dominance to improving the model’s reliability, tool integration, and user experience. Until then, the recommendation is to use Gemini 3.1 Pro for knowledge-intensive queries but rely on more consistent models like Claude or Codeex for coding and automation tasks. The video ends with a call for Google to make their models more usable, emphasizing that intelligence alone is not enough if the model cannot perform reliably in real-world scenarios.