Qwen 3.8 - Local AI xHigh Thinking is SCARY 🤯

The video showcases the Qwen 3.8 local AI model’s impressive ability to generate complex, high-fidelity content like photorealistic faces and detailed 3D models, while highlighting challenges such as thinking loops and runtime errors that arise during extensive token generation. To address these issues, the creator developed a thinking loop detector and experimented with various quantizations, temperature settings, and inference engines, positioning Qwen 3.8 as a powerful prototyping tool despite its current limitations.

In this video, the creator explores the capabilities of the Qwen 3.8 local AI model, focusing on its “extra high thinking” mode and extensive prompt testing. They experimented with various quantizations including Q4, Q8, Q9, and a new Q10, pushing the model to generate up to 140,000 tokens in some cases. A significant challenge encountered was the model’s tendency to enter thinking loops, repeatedly generating the same content. To address this, the creator developed a thinking loop detector that gracefully stops the generation when loops are detected, improving the overall output quality and efficiency.

The video showcases several impressive generative tasks, such as creating a photorealistic human face from scratch with 19,000 tokens, and a detailed 3D human anatomy model with organs and skeletal structure. Despite some minor visual glitches and runtime errors, these demonstrations highlight the model’s strong potential for complex, high-fidelity content generation. However, the model often produced runtime errors, especially when generating large amounts of code, which the creator suggests could be mitigated by running the model in iterative loops to fix errors progressively.

Various game-like environments were also tested, including Grand Theft Auto 5, Red Dead Redemption, Final Fantasy 7, and a 3D platformer similar to Super Mario. While the low thinking mode produced basic but playable versions, the extra high thinking mode often resulted in looping issues and runtime errors, requiring the thinking detector to manage. The creator noted that despite the errors, the model could generate impressive scenes, animations, and interactions, though it sometimes struggled with bugs like broken controls or missing elements.

The creator also experimented with temperature settings, finding that running the model at temperature zero yielded more deterministic and higher-quality outputs, while temperature one caused erratic looping behavior. Different inference engines were tested, revealing that looping issues persisted across them. The Q10 quantization, a higher fidelity setting, showed promise but still suffered from looping and errors, indicating that quantization level impacts stability and output quality.

Overall, the video presents Qwen 3.8 as a powerful 27 billion parameter local AI model capable of generating complex and creative content, but still prone to looping and runtime errors when pushed to its limits. The thinking loop detector is a key innovation to manage these issues. The creator emphasizes that while the model may not be perfect for final production, it serves as an excellent tool for prototyping and starting projects with rich UI and content generation. The extensive testing and token usage demonstrate the model’s potential and current limitations, inviting viewers to experiment and contribute to its development.

Useful Links