Continual Learning: How AI Agents Get Better With Every Use | Arjun Karanam, Trajectory

Arjun Karanam of Trajectory emphasized the importance of continual learning for AI agents, enabling them to improve through real user interactions by capturing detailed feedback and applying reinforcement learning techniques. Trajectory’s approach focuses on closing the experience gap in AI by ensuring traceability, realistic evaluation, effective tool harnesses, and flexible model management to create robust, adaptive AI systems that evolve with use.

Arjun Karanam, co-founder of Trajectory, presented on the topic of continual learning for AI agents, emphasizing the importance of experience alongside intelligence (IQ) in improving AI performance over time. He highlighted that while AI models are rapidly advancing in intelligence, they often behave as if it’s their “first day on the job,” lacking the accumulated experience that humans gain. Trajectory’s mission is to close this “experience gap” by enabling AI agents to learn continuously from real user interactions, thereby improving their capabilities as they are used.

The foundation of Trajectory’s approach starts with traceability—capturing detailed interaction data, including sub-agent actions and user feedback such as corrections and retries, which are often discarded. This data is then used to define precise model specifications that guide what the agent should learn. Their research focuses on converting user interactions into rewards and training signals, applying reinforcement learning techniques like SDPO to improve models over time. Additionally, they distinguish between knowledge that should be embedded in the model versus information better handled by external harnesses or tools, such as frequently changing facts.

Arjun outlined four key “wishes” or goals for advancing continual learning systems: first, comprehensive traceability that captures all interactions and elicits meaningful feedback; second, evaluation methods that closely mirror real user environments and workflows; third, harnesses designed not just to prevent errors but to empower agents to orchestrate product primitives effectively, with informative tool responses; and fourth, the ability to run and improve open-weight models, including experimenting with model routers to route tasks to the most suitable models. These wishes reflect the challenges and opportunities in building robust continual learning platforms.

During the Q&A, Arjun discussed the complexity of continual learning across different system components—models, harnesses, tools, and application layers—and emphasized the need for a holistic approach that optimizes learning across the entire system. He also addressed privacy concerns, noting techniques like synthetic data generation and distribution sampling to protect customer data while enabling learning. On episodic memory and feedback, he explained the importance of distinguishing between generalizable signals and user-specific preferences, training models accordingly while leaving some behaviors to context or harnesses.

Finally, Arjun shared insights on where continual learning is most impactful, particularly for frontier tasks where users push models beyond their current capabilities. He noted that continual learning enables models to improve in response to these challenging requests, expanding what users can achieve over time. Trajectory’s platform aims to empower companies to own their AI experience layers, making continual learning accessible and practical, with an easy-to-use interface for training, evaluating, and deploying improved models based on real-world interactions.