# Preprint finds mobile LLM inference less energy-efficient than batched servers

_Monday, September 14, 2026 at 12:00 AM EDT · AI, Science · Latest · Tier 2 — Notable_

A preprint studying 18 language-model configurations on two smartphones and a server found on-device inference was, on average, three times less energy-efficient than batched server inference. The authors measured energy per generated token, latency, accuracy and battery-cycle consumption. They report that 88% to 90% of modeled per-token environmental impact came from embodied device carbon rather than electricity use. The study also found non-monotonic energy effects from quantization and identified eight configurations on an accuracy-efficiency Pareto front.

## Sources

- [cs.LG updates on arXiv.org](https://arxiv.org/abs/2609.11940)

---
Canonical: https://techandbusiness.org/newswire/VAJuyBGCGy7HTCOLE-tthh
Retrieved: 2026-09-14T17:28:15.174Z
Publisher: Tech & Business (techandbusiness.org)
