Skip to main content
Back to Newswire
Science AI

Study measures intelligence per watt for local LLM inference versus cloud

Large language model queries still mostly run on frontier models in centralized cloud infrastructure, an arXiv preprint argues, even as demand growth strains that model. The authors propose intelligence per watt, task accuracy per unit of power, as a unified metric for capability and efficiency of local inference across model-accelerator setups. They evaluate more than 20 state-of-the-art local language models, eight hardware accelerators spanning local and cloud, and about 1 million real-world single-turn chat and reasoning queries. For each query they measure accuracy as local model win rate against frontier models, plus energy, latency, and power. Small local models at or below about 20 billion active parameters can now match frontier models on many tasks, and devices such as Apple's M4 Max can host them at interactive latencies, the paper says. Local models successfully answered 88.7% of the tested queries, with accuracy varying by domain. From 2023 to 2025, intelligence per watt improved 5.3 times, driven by algorithmic and accelerator advances, while locally serviceable query coverage rose from 23.2% to 71.3%. Local accelerators still achieved at least 1.4 times lower intelligence per watt than cloud accelerators running identical models, pointing to headroom for local hardware optimization. The authors conclude local inference can redistribute a substantial share of demand from centralized infrastructure for a large subset of queries.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.