Mahesh Sathiamoorthy of Bespoke Labs discussed the critical role of high-quality data and reinforcement learning environments in enhancing post-training large language models, highlighting their open-source tools like Curator and initiatives such as Open Thoughts for reasoning data curation. He emphasized that improving agent reliability through supervised fine-tuning and reinforcement learning, supported by curated data and environments, enables better long-term autonomy, reduced latency, and cost efficiency for enterprises, illustrated by practical applications like Credit Karma.
Mahesh Sathiamoorthy, co-founder and CEO of Bespoke Labs, presented on data and environment curation for post-training large language models (LLMs). Bespoke Labs focuses on providing enterprises and frontier labs with high-quality data and reinforcement learning (RL) environments to support post-training needs. Mahesh highlighted their open-source contributions, including the Curator tool for synthetic data curation, the Bespoke Stratos initiative for reasoning data, and the Open Thoughts project, which emerged from collaborations with academic institutions. Their work also extends to building and shipping RL environments and assisting enterprises in developing custom models through post-training.
Mahesh emphasized the evolution of AI evaluation from knowledge-based benchmarks to agent-based benchmarks that assess autonomous capabilities over extended durations. The key challenge in achieving long-term autonomy for agents is reliability, which can be improved through post-training techniques such as supervised fine-tuning (SFT) and reinforcement learning. While infrastructure and models are relatively mature, the bottleneck remains access to high-quality data and RL environments, especially for enterprises and frontier labs. Post-training not only enhances reliability but can also reduce latency and improve cost efficiency.
The Open Thoughts project was discussed in detail as an example of reasoning data curation. The team developed a systematic curation recipe involving selecting and mixing source questions, filtering, and generating answers using teacher models. They discovered that sampling multiple answers per question and using diverse reasoning traces improved model performance, while stronger teacher models were not always the best. This work demonstrated scalable improvements in reasoning benchmarks and gained recognition from industry leaders and researchers.
Building on this, Bespoke Labs extended their approach to agent training with Open Thoughts Agents, focusing on curating data and RL environments for agents. Similar lessons were learned, such as the benefits of multiple answer sampling and the surprising effectiveness of certain teacher models over others. They found that SFT contributed significantly to performance gains, with reinforcement learning providing incremental improvements at higher computational costs. A concrete enterprise example was shared involving Credit Karma, where post-training improved compliance, reduced latency, and enhanced throughput by curating data with specific tagging strategies to address data imbalance and hallucination issues.
Finally, Mahesh introduced the Curator tool, designed to simplify reasoning data curation by integrating with platforms like Tinker and Fireworks. He outlined a layered architecture for building RL models and post-training agents, encompassing environment creation, quality measurement, sandbox orchestration, and prompt optimization techniques such as Japa. This comprehensive stack supports scalable and efficient post-training workflows, aligning with emerging industry standards. Mahesh concluded by inviting further discussion and questions, highlighting the ongoing importance of data and environment curation in advancing autonomous AI systems.