# Nvidia adds cached inference workflow for sequential recommenders

_Published Wednesday, September 30, 2026 at 7:09 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![Nvidia adds cached inference workflow for sequential recommenders — Primary](https://developer-blogs.nvidia.com/wp-content/uploads/2026/09/recsys.webp)

Nvidia says Dynamo-Triton now supports an end-to-end workflow for serving HSTU recommendation models through its recsys-examples repository. These models rank items by processing sequences of user interactions. The workflow combines ahead-of-time PyTorch compilation with caching that reuses earlier sequence calculations instead of repeating them.

Nvidia reports best-case speedups of up to 4.47x for a three-layer model and 5.93x for an eight-layer model against the same compiled configuration without caching. Those results used batch size 8 on one RTX PRO 6000 Blackwell Workstation Edition GPU and required a 100% GPU cache hit rate.

## Sources

- [NVIDIA Technical Blog](https://developer.nvidia.com/blog/deploying-an-hstu-generative-recommender-with-nvidia-dynamo-triton/)

---
Canonical: https://techandbusiness.org/newswire/V7rCGNZkC95y7FWYGcb2iY
Published: 2026-09-30T23:09:00.114Z
Story chronology: 2026-09-30T20:54:59.000Z
Retrieved: 2026-10-01T01:18:49.584Z
Publisher: Tech & Business (techandbusiness.org)
