Science AI
arXiv study: semantic metadata still lifts FAIR precision for agentic data retrieval
An arXiv preprint in information retrieval asks whether large language model agents still need semantic metadata such as schema.org when retrieving machine-actionable datasets, or can navigate the open web alone.
The authors compare a Baseline Agent searching billions of open-web documents with a Semantic Agent using a corpus of 90 million datasets annotated with schema.org. An LLM-as-a-judge pipeline mapped to FAIR principles scores semantic relevance, accessibility, and computational utility of retrieved results.
The Semantic Agent achieved 44.9% higher precision for metadata-rich registries and 46.6% higher precision for pages with machine-readable downloads among its returned results. The Baseline Agent more often hit last-mile utility failures, including prose-heavy pages (20.1% of results) and portal landing pages (8.5%) rather than actual data pages.
The Baseline Agent answered 40% more questions and thus covered more queries, but the Semantic Agent delivered 65.7% higher overall precision on FAIR-compliant datasets. The authors conclude that unstructured retrieval supports broad exploratory tasks while structured metadata ecosystems remain foundational for reliable, execution-oriented autonomous workflows.
Sources
Published by Tech & Business, a media brand covering technology and business.
This story was sourced from arXiv and reviewed by the T&B editorial agent team.