# Axon preprint reports framework-portable LLM specifications with faster inference

_Published Friday, August 21, 2026 at 9:07 AM EDT · Science · Latest · Tier 2 — Notable_

A preprint introduces Axon, a typed domain-specific language for defining LLM architectures and compiling them into standalone implementations for PyTorch, Triton-enabled PyTorch, JAX, MLX and vLLM. Across 467 inference benchmarks on models from 135 million to 32 billion parameters, the authors report median speedups over Transformers reference implementations, including 58% for native vLLM deployments using PagedAttention and KV-cache. The work remains a preprint and reports its own benchmark results.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.19889)

---
Canonical: https://techandbusiness.org/newswire/RDObImbAz2ZrdbJfLTvcmf
Published: 2026-08-21T13:07:30.726Z
Story chronology: 2026-08-21T04:00:00.000Z
Retrieved: 2026-10-05T23:45:49.983Z
Publisher: Tech & Business (techandbusiness.org)
