# Amazon SageMaker adds automated concurrency tests for AI endpoints

_Published Tuesday, September 22, 2026 at 2:17 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![Amazon SageMaker adds automated concurrency tests for AI endpoints — Primary](https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21905-featured-image.png)

Amazon SageMaker AI now includes concurrency sweeps in Inference Recommendations, giving operators a managed way to measure how generative-AI endpoints behave as simultaneous requests increase. The benchmark records throughput and latency at each traffic level, exposing the point where throughput stops rising and requests begin to queue.

Operators can also use an SLA-based search to find the highest concurrency that remains within specified latency limits. AWS's example found 320 concurrent requests met a 50-second p99 latency target, but results depend on the model, hardware and workload profile tested.

## Sources

- [AWS Machine Learning Blog](https://aws.amazon.com/blogs/machine-learning/right-size-generative-ai-endpoints-with-concurrency-sweeps-on-amazon-sagemaker-ai/)

---
Canonical: https://techandbusiness.org/newswire/mJYjRuJP_Jc8ZzVzUAwTvO
Published: 2026-09-22T18:17:15.016Z
Story chronology: 2026-09-22T15:35:53.000Z
Retrieved: 2026-09-22T19:46:09.375Z
Publisher: Tech & Business (techandbusiness.org)
