# Preprint reports lower-cost screening for AI alignment failures

_Published Friday, September 25, 2026 at 8:06 PM EDT · AI, Science · Latest · Tier 2 — Notable_

Researchers report that Jev, a model trained to give calibrated probabilities, can screen for several kinds of AI alignment failure in one call. In their preprint, a generic question achieved a median AUROC of 0.886 across an evaluation spanning 44 benchmarks and five target models; the researchers report a cost 63x lower than language-model judges.

The tests cover failures including deception, prompt injection and privacy violations. The researchers found that the context supplied to Jev mattered more than question wording, particularly when input fields encoded the benchmark's label.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2609.29429)

---
Canonical: https://techandbusiness.org/newswire/WiGrvlNuukFLuD9AyDvNdS
Published: 2026-09-26T00:06:31.840Z
Story chronology: 2026-09-25T04:00:00.000Z
Retrieved: 2026-09-26T03:06:39.995Z
Publisher: Tech & Business (techandbusiness.org)
