# OpenAI details Jalapeño inference chip benchmarks and deployment timeline

_Tuesday, August 25, 2026 at 10:22 AM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![OpenAI details Jalapeño inference chip benchmarks and deployment timeline — Primary](https://techcrunch.com/wp-content/uploads/2026/08/Jalapeno-chip-final.jpeg?resize=1200,900)

OpenAI presented benchmark results for its Jalapeño inference system at the Hot Chips conference, saying SemiAnalysis' InferenceX test showed more tokens per user and more throughput per kilowatt than currently available state-of-the-art inference processors. OpenAI says the design keeps model state, including the KV cache, local to reduce data movement and communication delays during inference. The company expects very small deployment volumes at the end of 2026 and more significant deployment in 2027.

## Sources

- [techcrunch.com](https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/)

---
Canonical: https://techandbusiness.org/newswire/7-vwGxgjyAyaQ8w4mrQD0z
Retrieved: 2026-08-26T19:39:35.729Z
Publisher: Tech & Business (techandbusiness.org)
