text2sql-agent-eval: results explorer

100 questions over an Olist e-commerce warehouse, 5 difficulty tiers, 4 architectures. This is a replay of committed results; nothing runs live. Source code and build log.

Comparison

Single runs. Retry temperature (0.4) and hosted API calls both add run-to-run variance, so read these as accuracy ± a few points.

Question explorer

Tier:

Ground-truth SQL

Generated SQL

All four architectures on this question

ArchitectureResult
How correctness is judged

Correct means the result set matches the ground truth, not the SQL text. Column names and column order are ignored. Numbers match within 0.01 (or 0.01% for large values). Row order only matters when the question asks for an order. NULL matches NULL. So a query that looks different from the ground truth can still be marked correct.