Does Step 3.7 Flash Actually Beat DeepSeek & Claude? 🫣 200K TOKENS+

The video reviews Step 3.7 Flash from Step Fun AI, highlighting its impressive ability to generate nearly 200,000 tokens and excel in multimodal tasks like 3D scene creation and wiki management, though it struggles with accuracy in complex math problems and occasional runtime errors. While demonstrating strong creative and practical capabilities, the model shows mixed reliability and is not definitively superior to competitors like Deep Seek or Claude across all domains.

The video explores the capabilities of Step 3.7 Flash from Step Fun AI, a model touted to outperform competitors like Deep Seek V4 Flash and Claude CL Opus, especially in multimodal understanding and acting. The presenter highlights the model’s impressive token generation capacity, demonstrating a massive 198,471-token output that took over three hours to generate. Despite some initial confusion in finding the correct model online, the video dives into various tests including the Flappy Bird 3D challenge, Tiddly Wiki, car wash scenarios, and complex mathematical problems, showcasing both the strengths and limitations of the model.

In the Flappy Bird 3D challenge, the model performs well with different reasoning modes—low, medium, and high—producing visually appealing outputs and handling complex token generations, though some runtime errors and minor glitches occur. The cloud-based version runs faster but appears less reliable than the local runs. The Tiddly Wiki test demonstrates the model’s ability to manage and edit a local wiki effectively, with smooth token generation and functional tagging, indicating practical applications for documentation and note-taking.

The model’s performance on mathematical problems, however, reveals significant weaknesses. Despite generating nearly 200,000 tokens over three and a half hours, it fails to arrive at the correct answer for a challenging math Olympiad question, even with different reasoning modes enabled. This suggests that while the model can handle extensive reasoning tasks, accuracy in complex domains like mathematics remains an issue. The presenter notes that the model might be better suited for other domains such as legal or general knowledge, though caution is advised.

Graphical and 3D scene generation tests show promising results, with the model producing detailed SVG animations and 3D city scenes, albeit with occasional runtime errors and minor flaws like non-environmentally friendly lighting. The human face generation test also yields interesting results, producing faces with more features than some competitors, though higher reasoning modes sometimes fail. These tests highlight the model’s potential in creative and visual tasks, even if it struggles with consistency in more demanding reasoning scenarios.

Finally, the video tests Step 3.7 Flash’s integration with tool calls and the Open Claw evaluation framework, confirming its ability to access external tools and perform web requests effectively. The model successfully fetches live data, such as weather information, demonstrating practical utility in real-world applications. Overall, while Step 3.7 Flash shows impressive token handling, multimodal capabilities, and creative output, it remains a mixed bag in terms of accuracy and stability, positioning it as a strong contender but not yet a definitive leader over Deep Seek or Claude in all areas.