AWS launches GPU-aware SageMaker HyperPod Inference Gateway
AWS announced general availability of the SageMaker HyperPod Inference Gateway, a Kubernetes-native routing addon that places inference requests using real-time GPU signals such as KV cache utilization, queue depth, LoRA adapter r...
AWS ships next-gen AgentCore Runtime with elastic memory and 2-second cold starts
AWS announced general availability of a next-generation AgentCore Runtime, the serverless microVM compute layer inside Amazon Bedrock AgentCore. The company said the new runtime reclaims unused session memory instead of holding i...
Huawei outlines 2027 Ascend AI-chip launches
Huawei plans to launch its Ascend 960DT AI chip in the first quarter of 2027 and the Ascend 960PR in the third quarter, Reuters reported. Rotating chairman David Wang said UnifiedBus technology would be key to Huawei's next-genera...
PrismML releases 5.9 GB Bonsai 2 model compressed from Qwen 27B
PrismML on Thursday released Bonsai 2 27B, compressing Alibaba's open-source Qwen3.8 27B to 5.9 GB of memory, a 9x to 10x reduction the company says is small enough for a PC and possibly a high-end smartphone. The Caltech-founded...
GMI Cloud reportedly seeks $300 million chip loan for Thailand
GMI Cloud, an Nvidia partner, is seeking a $300 million loan to buy chips for a facility in Thailand, according to people familiar with the matter. The prospective financing would add to similar Asian deals supporting AI-compute i...
Manus reportedly targets $4 billion valuation after Meta split
Chinese-founded AI startup Manus is set to double its valuation to $4 billion in its first fundraising since Beijing ordered it to split from Meta Platforms, according to Bloomberg. The prospective round follows a geopolitical dis...
