RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

Brendan Foody from Mercor explains how advanced Reinforcement Learning environments, built with expert input and realistic workplace simulations, enable AI agents to learn complex real-world tasks by interacting with authentic data, software tools, and evaluative rubrics. He highlights the challenges of scaling expert involvement, ensuring data quality, and developing long-horizon, socially interactive tasks, while emphasizing the importance of specialized partners like Mercor in providing high-quality, scalable training data to drive frontier AI performance.

In this discussion, Brendan Foody from Mercor explains the evolution and significance of Reinforcement Learning (RL) environments in AI development, particularly focusing on how AI agents learn to perform real-world tasks. He traces the history from the early days of crowdsourced data for behavior cloning and supervised fine-tuning, such as RLHF, to the current shift towards an agentic data era. This new era emphasizes leveraging highly skilled experts across various professions—lawyers, doctors, engineers—to collaboratively build complex RL environments and frontier evaluations that push AI capabilities forward. Mercor has been instrumental in this transition, scaling up expert involvement and becoming a key data vendor for leading AI labs and application companies.

Foody outlines the structure of RL environments, which consist of three main components: worlds (realistic project data like messages and documents), apps (high-fidelity clones of popular software tools), and tasks (prompts and verifiers such as rubrics or unit tests). These environments simulate real workplace scenarios, enabling AI agents to learn tool use and task execution in a manner that mirrors human work. A significant challenge is covering the vast diversity of real-world jobs and applications, requiring massive expert input to ensure realism and accuracy. Humans remain essential, especially for creating verifiers that reliably assess AI performance, as models struggle to self-evaluate complex tasks without human-like feedback.

The conversation also touches on the economics and quality considerations of data used for RL training. Pricing is influenced by the value of model improvement to customers and the cost of expert labor, with tasks ranging widely in complexity and price. Data quality hinges on realism—how well environments reflect actual job conditions—and the accuracy of verifiers that score AI outputs. To maintain high standards, Mercor employs rigorous quality control processes, including trajectory analysis and human review, ensuring that automated scoring aligns with expert judgment. Synthetic data generation plays a complementary role, primarily through model-generated trajectories and assistance in environment creation, but human expertise remains crucial for frontier-level evaluation.

Looking ahead, Foody highlights emerging trends in RL environments, such as the need to handle ultra-long horizon tasks that span hundreds or thousands of hours and the introduction of virtual co-workers to better simulate social interactions in the workplace. These developments aim to close the realism gap in AI training by reflecting the collaborative and extended nature of many human jobs. He also discusses the bottleneck of rubric creation for task verification, noting that AI can assist experts but cannot yet fully automate this process due to the complexity and nuance involved in accurately grading AI performance.

Finally, Foody advises companies partnering with Mercor to balance leveraging custom, exclusive data sets that provide competitive advantage with the efficiencies gained from Mercor’s scalable infrastructure and talent network. He emphasizes that while some organizations attempt to build all capabilities in-house, the economies of scale and expertise offered by specialized partners are often critical to achieving frontier AI performance. The discussion concludes with insights into model limitations, the importance of base model strength for effective learning, and the ongoing evolution of RL environments as a foundational technology for AI applications across industries.