Thinking Machines, a $12 billion startup led by former OpenAI CTO Mira Murati, has launched Inkling, an open-weight, versatile AI model designed for efficiency and adaptability rather than top-tier performance, featuring capabilities like self-improvement and direct processing of raw audio and images. While Inkling’s performance is moderate compared to competitors, the company aims to offer it as a free, open mid-tier model and monetize through fine-tuning services, marking a significant milestone in their innovative approach to AI development.
Two years ago, Mira Murati, former CTO of OpenAI, made a bold move by quitting her job without a backup plan, driven by a desire for exploration. Unlike most people, her departure attracted significant attention and investment, with A16Z backing her new venture, Thinking Machines, to the tune of $2 billion. Despite being valued at $12 billion, the company had yet to ship a product—until recently. They launched Inkling, an open weights model capable of seeing, hearing, reasoning, fine-tuning, and potentially interacting with the real world, marking a significant milestone for the startup.
Mira’s team, which included key figures like co-founder John Schulman and VP of research Barrett Zoff, had previously released Tinker, an API for fine-tuning open-weight models without infrastructure hassles. While Tinker was met with moderate enthusiasm, Inkling represents a more ambitious leap. It is a massive mixture of experts model with 970 billion parameters, but cleverly activates only 41 billion per token, optimizing compute efficiency. Trained on 45 trillion tokens across text, images, and audio, Inkling supports a massive 1 million token context window and is openly licensed under Apache, making it accessible on Hugging Face.
Despite its impressive specs, Inkling’s performance is middling compared to competitors like Fable 5 and GPT 5.6 Sol, and it was quickly overshadowed by Moonshot’s Kimmy K3, a 2.8 trillion parameter model. However, Thinking Machines intentionally designed Inkling to be a versatile, cost-effective model rather than the outright smartest. It features a “thinking effort” dial that balances speed and accuracy, allowing users to tailor the model’s computational intensity to their needs. This flexibility is particularly valuable for applications requiring millions of agent runs daily, where efficiency translates to significant cost savings.
One of Inkling’s standout features is its ability to self-improve and self-regulate. In a demo, the model was tasked with removing the letter “E” from its outputs, then autonomously wrote its own training script, generated data, retrained itself, and updated its weights accordingly. Additionally, Inkling was trained on “epistemics,” enabling it to recognize and admit uncertainty rather than guessing confidently, making it one of the best models for forecasting future events. Unlike typical models, Inkling processes raw audio and pixels directly without separate encoders, though this approach leads to some quirks, such as a simplified, “caveman” style of internal language after extensive training.
Thinking Machines acknowledges that Inkling is not the top-performing model on the market, but that is by design. Their strategy is to provide a solid, open mid-tier model for free and monetize through Tinker’s fine-tuning capabilities, allowing users to create specialized models tailored to specific problems. The video also highlights Clerk, a sponsor offering tools to streamline billing and user management for developers, emphasizing the practical side of deploying AI-powered applications. Overall, Inkling represents a significant step for Thinking Machines, showcasing innovation in model design, openness, and practical usability.