{"version":"1.0","type":"rich","provider_name":"techandbusiness.org","provider_url":"https://techandbusiness.org","title":"AWS adds prefix-aware routing to SageMaker Inference for LLM endpoints","author_name":"techandbusiness.org · AI","thumbnail_url":"https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/10/ML-21885-featured-image.png","width":600,"height":400,"html":"<blockquote class=\"tb-newswire-embed\" style=\"max-width:600px;border-left:3px solid #22d3ee;padding:12px 16px;margin:0;font-family:-apple-system,system-ui,sans-serif;background:#09090b;border-radius:0 8px 8px 0;\">\n      <p style=\"margin:0 0 8px;font-size:10px;font-weight:600;letter-spacing:0.1em;color:#71717a;\">techandbusiness.org · AI</p>\n      <p style=\"margin:0 0 8px;font-size:18px;font-weight:700;line-height:1.3;color:#fff;\"><a href=\"https://techandbusiness.org/newswire/Vw2pjilvS3lkLPi8qW9e3g\" style=\"color:#fff;text-decoration:none;\">AWS adds prefix-aware routing to SageMaker Inference for LLM endpoints</a></p>\n      <p style=\"margin:0;font-size:14px;color:#a1a1aa;line-height:1.5;\">Amazon SageMaker Inference introduced prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so cached key-value pairs are reused instead of r...</p>\n    </blockquote>"}