# Preprint finds language models falter during sequential emergency triage

_Published Tuesday, September 22, 2026 at 12:09 PM EDT · Science, AI · Latest · Tier 2 — Notable_

A preprint evaluating six language models found that their emergency-department triage performance deteriorated when patient information arrived turn by turn rather than as a completed record. Across 425 simulated and 50 physician-authored conversations, the best model reached a quadratic weighted kappa score of 0.295, compared with 0.887 to 0.929 for three expert clinicians.

Controlled tests indicated that models anchored on the initial complaint despite extracting relevant information from later exchanges. Prompting did not correct the plateau, and combining models worsened results because their predictions agreed with one another more than with the ground truth.

## Sources

- [cs.CL updates on arXiv.org](https://arxiv.org/abs/2609.22904)

---
Canonical: https://techandbusiness.org/newswire/QqosH8e6UrvcpCZi7Pcx3-
Published: 2026-09-22T16:09:07.504Z
Story chronology: 2026-09-22T04:00:00.000Z
Retrieved: 2026-09-22T17:45:12.519Z
Publisher: Tech & Business (techandbusiness.org)
