Why Off-the-Shelf AI Doesn't Understand Money — Udi Menkes, Intuit

Udi Menkes explains that off-the-shelf AI models fail to provide reliable financial advice because they lack grounding in real-world business outcomes, often producing confident but misleading recommendations based solely on internet data. He advocates for AI systems trained on verified, outcome-based data from actual business experiences, emphasizing that true financial understanding comes from experience rather than just knowledge, which enables more trustworthy and effective advisory tools.

Udi Menkes begins by sharing a personal anecdote about his three-year-old daughter’s simplistic understanding of money, highlighting how even advanced AI models struggle with financial advice despite their fluency. He notes that while many people use large language models (LLMs) for financial recommendations, few fully trust or act on their advice. Menkes illustrates this with his own experience where slight changes in input assumptions led to drastically different investment recommendations, underscoring the unreliability of off-the-shelf AI in financial decision-making.

He presents real-world examples from Intuit’s research involving thousands of businesses, showing how leading frontier LLMs often give risky or harmful financial advice. For instance, a model suggested acquiring a second rental property to improve profit despite the business already being in negative cash flow, while a grounded model recommended a safer price increase based on actual outcomes from similar businesses. Another example involved an egg supplier where the frontier model advised a risky price hike, whereas the grounded model suggested negotiating vendor costs. Menkes terms this misleading confidence “the fluent bluff,” where AI confidently generates plausible but potentially damaging advice because it lacks grounding in real-world outcomes.

Menkes explains that the core issue is that frontier models “read about money” through internet data but have not “watched what happens” in real business scenarios. He contrasts this with grounded models trained on millions of business trajectories, combining state, action, and outcome data to learn which financial decisions actually lead to success. This approach, using reinforcement learning and real verified outcomes, allows Intuit to outperform larger frontier models with smaller, more focused models. The key advantage lies not in model size but in access to rich, outcome-based data.

He emphasizes the importance of experience over mere knowledge, comparing AI advisors to human advisers where experience and understanding of context matter more than textbook knowledge. Menkes highlights the challenge of measuring true impact in business decisions and advocates for AI systems that incorporate verified outcomes from similar businesses to provide trustworthy advice. He also stresses that great advisory experiences must combine solid scientific grounding with personalized understanding of the user’s preferences to build trust and engagement.

Finally, Menkes frames the future of AI as outcome-driven, where success depends on grounding models in unique, domain-specific data and verified results rather than relying solely on larger, generic models. He encourages AI and finance leaders to deeply analyze their data to identify patterns of actions and outcomes, embedding this experience into AI systems. This approach, he argues, will close the gap in AI’s understanding of money and enable transformative, reliable financial advisory tools. Menkes concludes by inviting the audience to explore their data post-holiday and build AI grounded in real-world outcomes.