Mistral 3.5 Medium BEATS Kimi AND Claude? 🤯 Local AI TEST & REVIEW

The video reviews the Mistral 3.5 Medium AI model, highlighting its impressive performance in coding and vision tasks despite its relatively smaller size, outperforming competitors like Claude and Kimiko 2.6 in benchmarks. While showcasing strengths in generating playable games and solving complex problems, the presenter also notes limitations in handling more complex 3D rendering, speed, and stability, concluding that the model holds strong potential but requires further community testing and improvements.

The video reviews the Mistral 3.5 Medium model, a 128 billion parameter AI model labeled as medium size, which is an upgrade from their previous small 119 billion parameter Mistral 4 model. Despite the confusing naming, Mistral 3.5 is the latest flagship model from the French AI company, touted for its strong performance in benchmarks, particularly in software engineering tasks. The presenter is impressed by the model’s ability to outperform competitors like Claude and Kimiko 2.6 in coding benchmarks, which is surprising given the model’s relatively smaller size compared to others like Quinn’s 397 billion parameter model.

The video also explores the model’s vision capabilities, showing that Mistral 3.5 can accurately identify images, such as distinguishing an orange tabby cat from a red fox, an improvement over the previous Mistral 4 small edition. The presenter experiments with quantization to reduce the model size for running on systems with limited memory, achieving functional versions down to 2.9 bits, which still recognize images and respond to prompts. However, more complex tasks like interpreting CT scans require higher bit precision, increasing the model size beyond typical consumer hardware limits.

When testing the model’s coding abilities, the presenter demonstrates that Mistral 3.5 can generate a playable Tetris game, which is impressive for a model released two years ago. However, attempts to create more complex games like Flappy Bird or 3D city scenes reveal limitations, including runtime errors and basic outputs. The model struggles with WebGL and 3D rendering tasks, producing simple or faulty results compared to other models like Qwen 27B or Kimiko 2.6, which generate more polished outputs. The presenter notes that enabling the model’s “high thinking mode” slows down generation and does not necessarily improve results.

In mathematical problem-solving tests, Mistral 3.5 performs well, correctly answering Olympiad-style questions and reasoning through problems with or without the high thinking mode enabled. The model also successfully retrieves web pages when prompted but fails to execute some local coding tasks. The presenter highlights some confusion around the model’s release, including issues with weight files and updates, and notes that speculative decoding, which could improve performance, is not yet implemented for this model.

Overall, the video presents Mistral 3.5 Medium as a promising and versatile AI model with strong benchmark results, especially in coding and vision tasks, but with some practical limitations in speed, complexity handling, and stability. The presenter questions the validity of some of the claimed benchmark superiority over larger models and invites viewers to share their experiences and use cases. The video concludes with a positive note on the model’s potential and a call for further exploration and community feedback.