DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
- Gokul Kumar
- Yotam Perlitz
- et al.
- 2026
- EMNLP 2026
Gokul Karthik Kumar is an AI Researcher & Engineer at IBM Research in Zurich and a Marie Skłodowska-Curie Actions PhD Fellow in the ARMADA project, co-funded by Switzerland’s SERI and the European Union. He is pursuing a PhD in Artificial Intelligence at TU Wien, with his research hosted at IBM. His work focuses on agentic code generation with LLMs. His first PhD project, DataKernelBench, evaluated LLMs on translating SQL into optimized CUDA and Triton kernels that outperform torch.compile on GPUs.
Previously, at the Technology Innovation Institute, he developed WavLink, an audio–text embedding model that achieved competitive performance with up to 8× lower-dimensional embeddings, and Falcon3-Audio, an audio LLM that outperformed prior open-weight models by more than 10 points across speech, music, and sound question-answering benchmarks. He also enhanced the distributed pretraining codebase for the Falcon3 LLM. At Microsoft Research and AI4Bharat, he released state-of-the-art text-to-speech models for 13 Indian languages. At MBZUAI, he developed Hate-CLIPper, a hateful-meme classifier that achieved an AUROC of 85.8, exceeding human performance. He also built AutoDub, a human-in-the-loop, AI-powered dubbing platform that received an award at the IEEE SLT International Hackathon and was featured in WIRED Middle East.
He holds an MSc in Computer Vision from MBZUAI and a BTech from Anna University. He has first-authored papers at EMNLP, ICASSP, ASRU, and PAKDD.
Personal website: gokul.ch.