Thinking Machines’ Inkling is the USA’s largest open-weight multimodal AI model, excelling in audio inference, tool-calling, and creative image and code generation despite operating under challenging low-bit quantization. Positioned as a promising open-source alternative with advanced safety and reasoning capabilities, Inkling showcases significant potential for future development and broader adoption in the AI community.
The video explores Thinking Machines’ Inkling model, touted as the USA’s largest open-weight AI model with nearly a trillion parameters. Founded by former OpenAI team members, Thinking Machines has released this general-purpose multimodal model capable of processing text, images, and notably audio, which sets it apart from many competitors. The presenter highlights its superior audio inferencing capabilities compared to other models like Nvidia’s Nematron 3 Ultra, despite running on a challenging three-bit quantization that typically reduces model fidelity. Inkling also features advanced tool-calling abilities, allowing it to autonomously search the web and cross-verify information from multiple sources, demonstrating impressive problem-solving and research skills.
Testing reveals that while Inkling struggles with complex mathematics under low-bit quantization, it excels in safety mechanisms and reasoning, avoiding harmful or illogical outputs. The model shows promise in generating code and handling multimodal inputs, such as creating 3D voxel art from images and producing coherent game simulations inspired by titles like Rocky Balboa, Grand Theft Auto, and Final Fantasy 7. Despite some graphical glitches and bugs attributed to quantization limitations, the model runs without runtime errors, which is a significant achievement at such low precision levels. The presenter notes that higher-bit quantization could unlock even greater potential.
Inkling’s image inference capabilities are demonstrated through creative outputs like photorealistic 3D faces and animated voxel cats, showcasing procedural texture generation without relying on external assets. The model’s ability to analyze and improve images based on user feedback further underscores its versatility. Although some generations appear surreal or abstract, these outputs reflect the model’s complex internal reasoning and creative potential. The presenter compares these results favorably to other contemporary models, emphasizing Inkling’s robustness despite its early development stage.
A standout feature of Inkling is its advanced audio inference, which accurately transcribes speech and identifies speaker characteristics such as gender and voice tone. This contrasts with other models that either misidentify these attributes or only provide basic transcription. The model’s ability to integrate audio understanding with text and image processing makes it uniquely powerful in multimodal AI applications. The presenter praises this capability as a major step forward for open-weight models, highlighting the relatively small size of the audio weights compared to the overall model.
In conclusion, the video positions Thinking Machines’ Inkling as a promising and commercially open AI model with significant potential for future development. Licensed under Apache 2, it invites broader use and experimentation. The presenter expresses excitement about upcoming versions that may scale to trillions of parameters and suggests collaborative approaches to overcome computational challenges. Overall, Inkling represents a noteworthy American contribution to the open-source AI landscape, combining multimodal processing, tool integration, and safety features in a large-scale model.