# Preprint reports lower training cost for audio-visual reasoning

_Published Monday, October 5, 2026 at 11:51 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A preprint reports that text-only post-training improved Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming a complete audio-visual training route while using 56.6% fewer GPU-hours.

The method combines supervised fine-tuning with reinforcement learning, which trains a model using feedback on its outputs. Text-only training degraded perception, so the researchers added a smaller audio-visual reinforcement-learning stage to refine that capability.

That refinement used about 90% fewer input tokens than full-data audio-visual reinforcement learning, restored perception above the base level and retained 93.5% of the best text-only pipeline's reasoning gain.

## Sources

- [cs.CL updates on arXiv.org](https://arxiv.org/abs/2610.02819)

---
Canonical: https://techandbusiness.org/newswire/lVzBc5euR5ZPJJeooD__ng
Published: 2026-10-05T15:51:05.943Z
Story chronology: 2026-10-05T04:00:00.000Z
Retrieved: 2026-10-05T18:10:27.975Z
Publisher: Tech & Business (techandbusiness.org)
