# Preprint describes million-scale LLM-native knowledge base

_Thursday, September 3, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

A new preprint describes GPTKB 2.0, a method for building a disambiguated knowledge base directly from large language models. The authors report a materialized database with more than 1 million disambiguated entities and 38.4 million triples, using on-the-fly disambiguation of entities, relations and classes. They characterize trade-offs among accuracy, scale and cost, and say the method and associated resource are available online. The work has not been independently validated.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.03729)

---
Canonical: https://techandbusiness.org/newswire/1VzSMC3fUIgFp_bW57MLmF
Retrieved: 2026-09-03T12:46:16.249Z
Publisher: Tech & Business (techandbusiness.org)
