# Preprint reduces unsupported conclusions by AI investigators through training

_Published Monday, October 5, 2026 at 9:20 AM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report in a preprint that training a 9B language model on investigation trajectories reduced overstatement of evidence from 97% to 35%. Correct conclusions that did not overstate the evidence rose from 3% to 43%.

The team built Nautil, a collection of 731 audited cases covering transport accidents, chemical safety, vehicle defects and production server incidents. The model requests evidence, revises hypotheses and decides whether to close an investigation or identify missing information. Removing the grounds for a conclusion lowered its closure rate by 26 points relative to a matched control.

Further reinforcement learning raised balanced closure accuracy from 69.2 to 83.3, but weakened how closely closure decisions depended on the evidence.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2610.03190)

---
Canonical: https://techandbusiness.org/newswire/BEIFbgDTUWMd1IjlcWoKXW
Published: 2026-10-05T13:20:40.263Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T14:47:39.268Z
Publisher: Tech & Business (techandbusiness.org)
