text2sql-agent-eval: results explorer
100 questions over an Olist e-commerce warehouse, 5 difficulty tiers, 4 architectures. This is a replay of committed results; nothing runs live. Source code and build log.
Comparison
Single runs. Retry temperature (0.4) and hosted API calls both add run-to-run variance, so read these as accuracy ± a few points.
Question explorer
Tier:
No question matches these filters.
Ground-truth SQL
Generated SQL
All four architectures on this question
| Architecture | Result |
|---|
How correctness is judged
Correct means the result set matches the ground truth, not the SQL text. Column names and column order are ignored. Numbers match within 0.01 (or 0.01% for large values). Row order only matters when the question asks for an order. NULL matches NULL. So a query that looks different from the ground truth can still be marked correct.