# Preprint finds LLM leaderboard ranks hinge on evaluation harness

_Tuesday, August 25, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

A preprint evaluating 12 open-weight instruction-tuned language models across 3,679 benchmark items says leaderboard outcomes can vary sharply with equally defensible evaluation-harness settings. Across 26 configurations, the authors report that gemma4-31b scored from 31% to 89%, while four of the 12 models reached first place under at least one configuration. The study says configuration-fragile items accounted for 95.7% of the average score gap between adjacent models and releases per-item records and analysis code.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.21382)

---
Canonical: https://techandbusiness.org/newswire/qG2OLs6X0D99Bly5Pz9lCq
Retrieved: 2026-08-25T08:43:17.045Z
Publisher: Tech & Business (techandbusiness.org)
