# Audit finds cheaper model can better judge text-to-SQL results

_Published Tuesday, September 29, 2026 at 8:07 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers auditing a production system that converts questions into database queries found that its deployed GPT-4o-mini evaluator frequently disagreed with human reviewers. On a set enriched for disputed cases, it flagged 77.1% of queries humans judged faithful.

A self-hosted Qwen3.6-27B evaluator achieved stronger agreement with human judgments and cost roughly 1/300 as much per call as Claude Opus 4.7, which scored similarly. The comparison between those two models covered 96 cases, limiting how firmly their agreement scores can be distinguished. The findings appear in a preprint.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2609.30290)

---
Canonical: https://techandbusiness.org/newswire/eVNt0eJuy3JMKRG8b8z20V
Published: 2026-09-29T12:07:54.463Z
Story chronology: 2026-09-29T04:00:00.000Z
Retrieved: 2026-09-29T14:09:49.041Z
Publisher: Tech & Business (techandbusiness.org)
