# Magnitude offers device-tuned inference for local coding agents

_Published Wednesday, September 30, 2026 at 4:21 PM EDT · AI · Latest · Tier 2 — Notable_

![Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on App... — Primary](https://opengraph.githubassets.com/bb595086ee0068d63573c89295057424108e4b96bdbcab9b7303f064bb73bf5f/magnitudedev/magnitude)

Magnitude offers an open source desktop inference engine that compiles and tunes model kernels on the user's device. Its developers claim 92% faster decoding on Metal and 19% on CUDA compared with llama.cpp, plus 27% less memory per agent. Concurrent sessions share cached prompt prefixes.

The app runs on macOS, Linux and Windows, using Apple Silicon, NVIDIA or AMD hardware, or a CPU. It connects to existing agents and exposes an OpenAI-compatible API. Models and prompts stay local; supported model families depend on its optimized kernels.

## Sources

- [github.com](https://github.com/magnitudedev/magnitude)

---
Canonical: https://techandbusiness.org/newswire/HpB_Js2CzR8dzSghGo2X1Z
Published: 2026-09-30T20:21:04.803Z
Story chronology: 2026-09-30T17:37:40.000Z
Retrieved: 2026-09-30T22:20:02.383Z
Publisher: Tech & Business (techandbusiness.org)
