# DwarfStar 4 offers local inference through selective model compression

_Published Friday, October 2, 2026 at 4:26 PM EDT · AI · Latest · Tier 2 — Notable_

![DwarfStar 4 logo and local inference project card — Primary](https://dwarfstar.sh/og-preview-v2.png)

DwarfStar 4 offers a C inference engine for running supported DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next models on high-memory Mac, CUDA and ROCm machines, its project site says. The MIT-licensed software combines text and vision models, local APIs, a command-line interface and a native agent.

The engine compresses routed model experts to two bits while preserving precision in critical shared paths. It also saves long prompt prefixes to SSD for reuse after restarts. The project lists V4 Flash Q2 as its baseline and says V4.1 Q2 streams from SSD; its published performance figures are estimates from its benchmark table.

## Sources

- [dwarfstar.sh](https://dwarfstar.sh/)

---
Canonical: https://techandbusiness.org/newswire/IHuFR2UFLLTSb3RAZQW7l0
Published: 2026-10-02T20:26:33.923Z
Story chronology: 2026-10-02T18:01:16.000Z
Retrieved: 2026-10-02T22:57:03.710Z
Publisher: Tech & Business (techandbusiness.org)
