Anthropic’s Sonnet 5 was tested on building a browser-based Age of Empires clone by first creating its own game engine and toolset, successfully producing a functional game with effective asset management, though its output was less polished than Opus 4.8. While Sonnet 5 offers cost advantages and competence, Opus 4.8 delivers superior visuals, smoother gameplay, and more dynamic features, highlighting the benefit of agents developing their own tools to overcome development challenges.
Anthropic recently released Sonnet 5, which reportedly performs similarly to Opus 4.8 on the GENTI Coding task but at a lower cost. To verify these claims, the creator tested Sonnet 5 by asking it to build a browser-based Age of Empires clone, with a unique challenge: the model had to first develop its own game engine and toolset to create and test the game. This approach was inspired by previous experiences where agents struggled with game development due to limitations like frame rates and visual inconsistencies. By building its own tools, the agent could better manage these challenges and improve the quality of the final product.
The agent developed several tools, including an asset viewer and a character studio, which allowed it to inspect and fix animations and assets frame-by-frame. This was crucial because the agent could not interact with the game in real time and had to rely on screenshots, which could be outdated by the time they were analyzed. By slowing down animations and creating a controlled environment, the agent could effectively debug and refine the game assets. This method not only improved game development but also has broader applications, such as building testing frameworks for enterprise software.
After about two hours, Sonnet 5 successfully produced a working Age of Empires-style game with functional characters, resource harvesting, and map exploration. The toolset it created enabled detailed asset management and animation control, resulting in a visually coherent and playable game. The creator also ran the same prompt with Opus 4.8 for comparison, using identical instructions and corrective prompts to ensure a fair evaluation.
Opus 4.8 delivered a noticeably better game, with improved visuals, smoother character interactions, and more dynamic gameplay, including combat and AI behavior. Its toolset was also impressive, featuring an asset studio and animation lab that facilitated detailed asset and animation management. The creator felt that Opus 4.8’s output was superior in quality, despite similar time and cost estimates for both models. This comparison highlighted that while Sonnet 5 is competent and cost-effective, Opus 4.8 currently produces more polished results.
In conclusion, the video demonstrates the value of having coding agents build their own tools to overcome environmental limitations and improve output quality. While Sonnet 5 shows promise and cost advantages, Opus 4.8 still leads in delivering higher-quality game development results. The creator invites viewers to share their opinions and promotes his Agenic Labs community and masterclass, which teach how to effectively use coding agents for real-world software projects.