# Axon preprint reports framework-portable LLM specifications with faster inference

_Friday, August 21, 2026 at 12:00 AM EDT · Science · Latest · Tier 2 — Notable_

A preprint introduces Axon, a typed domain-specific language for defining LLM architectures and compiling them into standalone implementations for PyTorch, Triton-enabled PyTorch, JAX, MLX and vLLM. Across 467 inference benchmarks on models from 135 million to 32 billion parameters, the authors report median speedups over Transformers reference implementations, including 58% for native vLLM deployments using PagedAttention and KV-cache. The work remains a preprint and reports its own benchmark results.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2608.19889)

---
Canonical: https://techandbusiness.org/newswire/RDObImbAz2ZrdbJfLTvcmf
Retrieved: 2026-08-21T14:06:40.928Z
Publisher: Tech & Business (techandbusiness.org)
