In a recent livestream, the speaker expressed disappointment with OpenAI’s GPT-4.5, comparing it unfavorably to the final season of “Game of Thrones” due to minimal improvements and underperformance against competitors like DeepSeek V3. They criticized the high API pricing and lack of groundbreaking features, suggesting that OpenAI has relied on scaling rather than innovation, while reserving final judgment until they can test the model themselves.
In a recent livestream discussing the announcement of OpenAI’s GPT-4.5, the speaker expressed significant disappointment with the new model, likening the experience to the infamous final season of “Game of Thrones.” Despite the anticipation surrounding GPT-4.5, the speaker felt that the improvements were minimal, primarily noting that it “hallucinates less” and feels more natural. They questioned the rationale behind comparing GPT-4.5 only to GPT-4.0, suggesting that the expectations were not met given OpenAI’s previous reputation for innovation in AI models.
The speaker highlighted that GPT-4.5 underperformed against DeepSeek V3, an open-source model released just two months prior. They pointed out that while GPT-4.5 excelled in some areas, such as science and multilingual benchmarks, it lagged behind DeepSeek V3 in math and coding benchmarks. This performance disparity was particularly striking given the resources OpenAI invested in developing GPT-4.5 compared to the relatively low cost of running benchmarks for DeepSeek V3.
The discussion also touched on the financial implications of using GPT-4.5, noting that its API pricing was exorbitantly high compared to DeepSeek V3, despite similar performance levels. The speaker criticized OpenAI for not delivering a more competitive product, especially considering the advancements made by other models in the market. They expressed frustration over the lack of groundbreaking innovations from OpenAI, suggesting that the company has been relying on scaling rather than introducing revolutionary changes.
The speaker further lamented the absence of impressive features that were once hallmarks of OpenAI’s models, such as advanced math capabilities and engaging demos. They questioned the quality of the training data used for GPT-4.5, implying that it fell short of expectations given OpenAI’s resources and expertise. The emphasis on pre-training was noted, but the speaker felt that it did not translate into the expected performance improvements.
In conclusion, the speaker indicated that they would reserve final judgment on GPT-4.5 until they had the opportunity to test it themselves, as they were unwilling to pay a high subscription fee for early access. They also mentioned plans to review other models, such as Clot 3.7 and Grock 3, in the meantime. The video ended with a call to action for viewers to subscribe to their newsletter for updates on cutting-edge AI research and a shoutout to supporters on Patreon and YouTube.