QWEN 3.8 27B Local AI Review

The video reviews the QWEN 3.8 27B AI model running locally via VLLM, showcasing its impressive zero-shot ability to generate high-quality, playable arcade-style retro games with functional sound and leaderboards, marking a significant improvement over previous versions. The reviewer highlights the model’s enhanced tool integration, efficient multi-GPU performance on a powerful hardware setup, and smooth workflow orchestrated by the Hermes agent, concluding that QWEN 3.8 represents a major advancement in local AI code generation for complex development tasks.

The video presents a detailed review of the QWEN 3.8 27B AI model running locally via VLLM, highlighting the first outputs generated after a 2 hour and 40-minute run. The user tasked the model with creating an arcade-style retro game emulator featuring three games with 8-bit synthwave aesthetics and sound, all coded in JavaScript without Python involvement. The games, including Void Raiders, Nebula Drift, and Starbreaker, were playable and featured functional sound and leaderboards, showcasing impressive zero-shot generation quality. Despite minor bugs such as imperfect sound toggling and control quirks, the overall gameplay experience was enjoyable and demonstrated a significant leap in code quality compared to previous versions like QWEN 3.6.

The reviewer emphasized the substantial improvements in tool calling and terminal operations, noting zero failures during the run, which contrasts with earlier experiences using QWEN 3.6 that required iterative fixes and struggled with certain functionalities like Playwright integration. The fresh Hermes agent installation used to orchestrate the process further contributed to the smooth workflow. The model effectively utilized its vision capabilities as instructed, and the reviewer praised the quality of the generated code as being far superior to prior iterations, especially given that this was achieved in a zero-shot setting without iterative prompting.

Performance-wise, the setup ran on a powerful WRX80 Threadripper platform equipped with four RTX 3090s and one RTX 4090, delivering noticeable speed improvements and efficient GPU memory utilization. The reviewer shared insights into the optimized run block configuration for VLLM, recommending disabling speculative config and enabling thinking mode, which was previously disabled in QWEN 3.6 to avoid workflow issues. The system achieved high token throughput rates and stable memory usage, with detailed GPU management and threading settings tailored for optimal performance in multi-GPU environments.

The video also touched on hardware considerations, including the thermal challenges posed by NVLink configurations and the potential need to reconsider its use due to heat buildup. The reviewer mentioned ongoing testing and plans to refine the setup further. Additionally, they highlighted the importance of using internal Git repositories like Giddy for managing agent skills and code commits, which enhances version control and rollback capabilities. The reviewer expressed enthusiasm for the model’s capabilities and the overall system build, promising more content and updates in the future.

In conclusion, the QWEN 3.8 27B model represents a significant advancement in local AI code generation, delivering high-quality, playable game code with minimal errors and excellent tool integration. The combination of a robust hardware setup, optimized software configuration, and the Hermes agent orchestration resulted in a smooth and efficient workflow. The reviewer encouraged viewers to explore the provided resources, including prompts and playable demos, and thanked supporters for their contributions. The video serves as both a technical overview and an endorsement of the QWEN 3.8 model’s potential for complex AI-driven development tasks.