# Preprint releases bibliographic benchmark with 62,899 synthetic documents

_Friday, September 4, 2026 at 12:00 AM EDT · Science · Latest · Tier 2 — Notable_

A preprint introduced SHELF, a synthetic benchmarking system for bibliographic work, and released its source code and data under permissive licenses. Its first release contains 62,899 model-written documents derived from Library of Congress vocabularies and supports classification, clustering, retrieval, pair classification and instruction retrieval tasks. The authors compared sparse methods, encoders and zero-shot decoders where applicable. They report that SHELF scores do not estimate accuracy on production catalogue data.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.03047)

---
Canonical: https://techandbusiness.org/newswire/mLYD7Jq4a1-96Yf3QpuQaE
Retrieved: 2026-09-04T22:37:02.423Z
Publisher: Tech & Business (techandbusiness.org)
