AI Science
Preprint finds LLM leaderboard ranks hinge on evaluation harness
A preprint evaluating 12 open-weight instruction-tuned language models across 3,679 benchmark items says leaderboard outcomes can vary sharply with equally defensible evaluation-harness settings. Across 26 configurations, the authors report that gemma4-31b scored from 31% to 89%, while four of the 12 models reached first place under at least one configuration. The study says configuration-fragile items accounted for 95.7% of the average score gap between adjacent models and releases per-item records and analysis code.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire