How Web Data Infrastructure Powers the Next Generation of AI — Patricija Žemaitytė, Oxylabs

Patricija Žemaitytė of Oxylabs highlights the vital role of scalable, low-latency web data infrastructure in enabling next-generation AI to access fresh, real-time information, emphasizing that innovation stems from rapid adaptation to evolving client needs and web environments. Oxylabs’ experience in building robust systems capable of handling massive request volumes demonstrates that effective AI depends not only on advanced models but also on reliable infrastructure that seamlessly integrates live web data into AI workflows.

Patricija Žemaitytė from Oxylabs discusses the critical role of web data infrastructure in powering the next generation of AI. She begins by emphasizing that while AI models are often the focus, the underlying infrastructure that provides fresh, real-time, and usable data is equally essential. Oxylabs, founded in 2015, specializes in building scalable infrastructure that enables companies to extract public web data efficiently. As AI shifts from relying solely on static training data to integrating live external information, this infrastructure becomes indispensable for keeping models relevant and effective.

Patricija shares a compelling story about how Oxylabs rapidly developed a video API for AI training within two weeks, despite the massive scale and complexity involved. Initially starting as a simple downloader, the product evolved through continuous adaptation to include support for transcripts, subtitles, metadata, and channel information. This iterative process highlighted a key lesson: innovation is driven by the ability to adapt quickly under pressure rather than following a fixed roadmap. The evolving product suite eventually became a comprehensive solution that met the diverse and growing needs of clients.

The discussion then shifts to the importance of speed and latency in AI data delivery, particularly with search engine results pages (SERP) data. Traditional scrapers had latencies around four seconds, which was too slow for AI applications requiring real-time interaction. Oxylabs tackled this challenge by redesigning their systems from scratch to achieve sub-second latency, ultimately delivering results in around 550 milliseconds. This breakthrough enabled AI workflows to interact with fresh data seamlessly, demonstrating that speed is not just a performance metric but a fundamental product requirement in the AI era.

Scaling infrastructure to handle massive volumes of requests presents another significant challenge. Oxylabs faced the daunting task of increasing their capacity from 10,000 to 60,000 requests per second in under two months, involving complex operations like routing, rendering, proxy handling, and data normalization. Achieving this scale required robust architecture, reliable central components, and sophisticated observability tools to monitor system health. The team learned that real-world production traffic testing was crucial, as synthetic tests could not fully replicate the complexities of actual client usage.

In conclusion, Patricija underscores that Oxylabs is much more than a proxy provider; it is a builder of adaptable, scalable infrastructure that connects AI models to the dynamic web environment. The ongoing challenge is continuous adaptation to changing web layouts, detection methods, and client needs, making innovation an endless process. The future of AI depends not just on better models but on better infrastructure that bridges models to reality, enabling seamless integration of live web data into AI pipelines and workflows. This infrastructure foundation allows AI companies to focus on building intelligence while Oxylabs manages the complex maintenance and scaling demands.