Stop Overpaying for Intelligence | DevDay 2026

The DevDay 2026 panel emphasized optimizing AI costs by focusing on cost per task rather than cost per token, advocating for selecting models based on task complexity and accuracy needs while leveraging strategies like prompt caching, programmatic tool calling, and adjustable reasoning effort. They highlighted the importance of iterative experimentation and using OpenAI’s tools to balance cost and performance effectively, encouraging a holistic approach beyond simply choosing the cheapest model.

The panel discussion at DevDay 2026 focused on optimizing AI costs, emphasizing the importance of measuring cost per task rather than cost per token. Mandeep, leading the applied AI engineering team, explained that a cheaper model per token might not be cost-effective if it requires more tokens or human intervention to complete a task. He illustrated this with a personal story about a customer support chatbot failing to reschedule his flight, which ultimately required human assistance, increasing the overall cost. Real-world examples from customers like Perplexity and Notion demonstrated that investing in more capable models can reduce cost per task by improving accuracy and efficiency.

Choosing the right model size and configuration is crucial for cost optimization. Mandeep advised developers to clearly define the task and the accuracy required before selecting a model. He introduced the concept of the Pareto curve to balance accuracy and cost, encouraging experimentation with different models and settings to find the most cost-effective solution. For simpler tasks, smaller models like Luna might suffice, while more complex tasks benefit from advanced models like 6 Astra. Luna maxing, a technique where Luna is used extensively before escalating to larger models, was highlighted as an effective strategy for many workflows.

Beyond model selection, Sapto discussed four additional levers to optimize costs: prompt caching, programmatic tool calling, reasoning effort, and batch or flex API processing. Prompt caching reuses repeated input text to reduce processing costs significantly. Programmatic tool calling offloads data processing to code rather than the model, reducing token usage and model calls. Adjusting reasoning effort allows developers to balance the depth of model processing with cost, starting with lower effort for routine tasks. Batch or flex API processing offers lower pricing for non-urgent requests by allowing delayed processing.

The discussion delved deeper into prompt caching, with Sapto emphasizing common mistakes such as inconsistent prompt beginnings that reduce cache hit rates. He explained new features like explicit breakpoints and explicit-only mode in the latest models, which give developers more control over what parts of the prompt are cached. These improvements help maximize cache reuse, significantly lowering costs and latency. Mandeep added that developers can now change reasoning effort and tools without breaking the cache, enhancing flexibility. OpenAI also provides tools like a prompt caching dashboard and diagnostics API to help developers monitor and optimize cache performance.

In conclusion, the speakers urged developers to start by defining their tasks and accuracy goals, then experiment with different models and settings to find the optimal balance of cost and performance. They recommended iterative testing of changes, leveraging prompt caching and other levers to reduce costs while maintaining quality. The session highlighted OpenAI’s commitment to providing models at the Pareto frontier of cost and accuracy and encouraged attendees to engage with the AI deployment engineering team for further guidance. The overall message was clear: optimizing AI costs requires a holistic approach beyond just picking the cheapest model.