Autoresearch Made Our Models 3x Faster — Tejas Bhakta, Morph

Tejas Bhakta discusses how Auto Research, a framework that combines human insight with automated parameter tuning, can optimize low-level GPU operations to achieve up to three times faster model inference by iteratively testing and adjusting kernel parameters while maintaining correctness. He emphasizes that while Auto Research significantly accelerates GPU models, it requires detailed hardware knowledge, careful reward design, and substantial human guidance to overcome challenges and realize its full potential.

Tejas Bhakta explains how Auto Research, a framework developed by Andriy Karpaty, can be used to speed up GPU models by three times. Auto Research works by creating an environment where an agent iteratively tests and adjusts parameters to optimize performance towards a defined goal, such as increasing speed while maintaining correctness. This process is particularly effective for tuning low-level GPU operations, or kernels, which are fundamental for parallel computations on NVIDIA GPUs. While Auto Research excels at fine-tuning parameters like block sizes, it requires human input for higher-level architectural decisions.

The key to successful optimization lies in combining human insight with Auto Research’s parameter selection. Humans provide the innovative ideas and general direction, while the agent handles the detailed tuning and performance validation. Profiling tools like Nvidia’s Nsight help identify bottlenecks such as compute or memory inefficiencies, which guide the optimization process. Tejas emphasizes that Auto Research is especially useful for optimizing cheaper GPUs that lack ready-made cores, but this requires building a dedicated testing and auto-research framework tailored to the hardware.

For the agent to work effectively, it must have detailed knowledge of the hardware specifics, such as warps, TMEM, and TMA, which vary between GPU generations. Additionally, understanding the model’s unique features, like new attention mechanisms in DeepSeek Flash, is crucial to avoid generating ineffective kernels. A major challenge is designing a reward system that discourages the agent from making changes that degrade overall performance, such as disabling CUDA graphs or testing only on small context windows, which can lead to misleading speed improvements.

Tejas also notes that not all optimized kernels will perform well across all scenarios; some may only be effective within certain context sizes, requiring fallback to standard kernels in other cases. However, the cumulative effect of multiple optimized kernels can push GPU utilization closer to its theoretical maximum, especially when combined with hardware-level tweaks like BIOS adjustments, overclocking, or PCIe relaxing mode. These combined software and hardware optimizations can yield up to a threefold speed increase in model inference.

In conclusion, while Auto Research can significantly accelerate GPU models, it is not a magic solution and requires substantial human guidance and experimentation. Most attempts will fail, but persistence can lead to impressive results. Tejas encourages having original ideas as the foundation for optimization and invites interested individuals to join their efforts. The overall message is that the synergy between human creativity and automated parameter tuning is key to achieving breakthrough performance improvements.

Useful Links