# Nvidia reports confidential AI inference retains more than 96% of throughput

_Published Tuesday, September 22, 2026 at 3:08 PM EDT · Infrastructure, AI · Latest · Tier 2 — Notable_

![Nvidia reports confidential AI inference retains more than 96% of throughput — Primary](https://developer-blogs.nvidia.com/wp-content/uploads/2024/12/cybersecurity-graphic-1.png)

Nvidia says confidential DeepSeek-R1 inference retained more than 96% of ordinary output-token throughput on eight B200 GPUs after TensorRT-LLM was adapted for secure execution. Across concurrency levels from one to 16, confidential computing retained 96.1% to 98.2% of baseline throughput, while mean time per output token rose 1.2% to 4.3%.

The framework changes how protected host-to-device copies, timing measurements and multi-GPU communications are handled. The controlled test used a long-context, extended-generation workload at low concurrency to expose overhead; results may differ for other models, traffic patterns and configurations.

## Sources

- [NVIDIA Technical Blog](https://developer.nvidia.com/blog/enabling-private-high-performance-production-ai-inference-with-nvidia-confidential-computing/)

---
Canonical: https://techandbusiness.org/newswire/BvX5SO6koH9JgbUlJxheiL
Published: 2026-09-22T19:08:10.728Z
Story chronology: 2026-09-22T17:27:41.000Z
Retrieved: 2026-09-22T20:41:14.164Z
Publisher: Tech & Business (techandbusiness.org)
