# Preprint describes VLM inference method with reported 3.1x faster first token

_Friday, August 28, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

An arXiv preprint describes PACE, a training-free framework that combines input downsampling before vision encoding with selective visual-token retention after encoding. Integrated into Qwen2.5-VL-7B, its authors report retaining 93.8% of the model's original performance while using 10% of visual tokens and achieving a 3.1x time-to-first-token speedup. The paper says code is available. These are author-reported experimental results, not an announced product release.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.27206)

---
Canonical: https://techandbusiness.org/newswire/JfRVmDsC6UGq0esDkgWkZg
Retrieved: 2026-08-28T09:49:58.730Z
Publisher: Tech & Business (techandbusiness.org)
