Manuel Le Gallo, Corey Liam Lammie, et al.
APL Mach. Learn.
GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves 2.11x speedup over torch.compile at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving 2.54x speedup.
Manuel Le Gallo, Corey Liam Lammie, et al.
APL Mach. Learn.
Corey Liam Lammie, Hadjer Benmeziane, et al.
SOCC 2026
Sahil Suneja, Yufan Zhuang, et al.
ACM TOSEM
Chih-kai Ting, Karl Munson, et al.
AAAI 2023