What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

Nebius Token Factory is a comprehensive platform that optimizes the deployment and management of open-source large language models by combining managed inference services, advanced caching, quantization, and fine-tuning tools on their own bare-metal infrastructure to deliver fast, cost-effective, and customizable AI solutions. By addressing challenges in load balancing and resource optimization, Nebius enables enterprises to efficiently run and continuously improve LLMs at scale without the complexity of self-hosting or the limitations of closed APIs.

In this presentation, Dylan and Sujee from Nebius introduce their product, Nebius Token Factory, which is designed to optimize the deployment and management of open-source large language models (LLMs) in production environments. Nebius operates its own data centers with bare-metal infrastructure, providing a vertically integrated AI cloud stack that offers customers control, performance, and cost benefits. They emphasize the challenges AI teams face when choosing between closed APIs, which are easy but limited and costly, and self-hosting, which is complex and resource-intensive. Token Factory aims to bridge this gap by offering managed inference services that combine the simplicity of APIs with the control and customization of self-hosting.

The platform provides a full-stack solution that integrates model inference, data logging, post-training fine-tuning, and deployment into a continuous improvement loop. Customers can run over 60 models with various optimizations such as structured outputs and function calling. The Data Lab component allows users to capture and analyze production logs, enabling better understanding and refinement of model behavior. Post-training tools support fine-tuning, distillation, and quantization to tailor models to specific use cases. Finally, the deployment layer ensures smooth updates and control over model versions and infrastructure, all managed on Nebius’s own hardware.

Sujee highlights the competitive performance of open-source models compared to proprietary ones, noting that the gap in intelligence and cost-efficiency is narrowing. Nebius hosts a range of large “teacher” models and smaller “student” models, optimizing deployment by selecting the best serving engines for each model. They address the complexity of load balancing in LLM inference by using cache-aware routing to improve efficiency and speed. This approach ensures that requests are directed to GPUs with relevant cached data, significantly enhancing performance.

Several technical optimizations underpin the platform’s speed and cost-effectiveness. Spec decoding uses smaller, faster models to generate tokens, with larger models verifying outputs to balance speed and accuracy. KV caching stores previously generated tokens to avoid redundant computation, yielding speed improvements of up to 10x. Nebius also separates the prefill and decoding stages of inference across different GPUs to optimize resource usage. Additionally, they apply careful quantization to reduce model precision without sacrificing quality, further improving efficiency.

The presentation concludes with an invitation to try Nebius Token Factory, which supports the latest open-source models and offers features tailored for coding and other applications. The team encourages engagement through their Discord channel and direct contact for questions. Overall, Nebius positions Token Factory as a comprehensive, enterprise-ready platform that simplifies running open-source LLMs at scale, enabling teams to focus on building AI products rather than managing complex infrastructure.

Useful Links