New research from Turing, in collaboration with Oracle: PLSQLBench, a benchmark for evaluating whether LLMs can write executable PL/SQL programs.
Most benchmarks test general code generation or text-to-SQL. But real database work isn't a one-shot query, it's stored procedures, cursors, exception handling, and iterating against a live schema.
That entire layer has gone basically untested. Until now.
Introducing PLSQLBench: the first benchmark built for procedural database programming.
-2,865 instances
-2,594 single-turn tasks + 271 multi-turn conversations (978 turns)
Built on enterprise-style Spider 2 databases, schema-grounded Spider tasks, and MBPP-derived procedural problems.
It measures what developers actually do: write procedures, handle exceptions, ground work in real schemas, and refine iteratively.
If we want LLMs that ship real database code, we have to evaluate them on real database work.
Testing eight LLMs surfaced recurring weaknesses in schema grounding, PL/SQL dialect fidelity, procedural control flow, exception handling, and consistency across turns. Tool-augmented agents closed some of the gap on schema-grounded tasks, but meaningful gaps remain.
The paper has been accepted to EMNLP industry track 2026 in Budapest.
Paper + Code below.
Congratulations to all: Marianne Menglin Liu, Leo (Leonid) Boytsov,
Daniel Petersen, Pramuditha Perera, Rongguang Wang, Sai Ashish Somayajula, Syed Hamza Rafique, Rohit Saini, Shubham Pathak, Sujeeth Bharadwaj, Tao Sheng, Graham Horwood, Fahad Shah, Ankan Bansal, Sujith Ravi, Dan Roth