Skip to main content
Infrastructure

Smaller, faster, safer: running Kimi and GLM at scale

Smaller, faster, safer: running Kimi and GLM at scale Image: Primary
Cloudflare said it has deployed three optimization techniques to run Moonshot's Kimi K-series and Z.ai's GLM mixture-of-experts models on its Workers AI platform. The company quantizes the key-value cache to 8-bit floating point, compresses model weights to 4-bit integers, and adds integrity checks to protect the shared cache. Quantizing the KV cache on Kimi K2.6 doubles the context capacity to about 1.37 million tokens and allows 64 concurrent requests instead of 32. Cloudflare said this reaches 2,192 tokens per second, about 41% higher than the 16-bit precision peak, for roughly 30% less cost per token. Compressing GLM 5.2 weights from 8-bit to 4-bit integers shrinks the checkpoint from 705 GB to 421 GB and reduces per-GPU memory from 88 GB to 52 GB. The company said decode speed improves because less data moves across memory bandwidth, while prefill remains in higher precision to avoid a slowdown. Cloudflare added a KV cache integrity checking layer that tags physical cache pages and validates mappings before decode operations read from the cache. The company said the safety check costs under 1% on both throughput and tail latency in production measurements. All work runs on the SGLang inference serving framework.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from blog.cloudflare.com and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Security Infrastructure
Security Infrastructure

Port of Los Angeles reports 120 million cyberattack attempts

The Port of Los Angeles foiled more than 120 million cyberattack attempts in August, according to Bloomberg. The U.S.'s busiest container port for global trade faces a persistent operational threat while navigating shifting tariff...

Infrastructure
Infrastructure

Amazon Leo adds six Ariane 6 launch flights

Amazon said Amazon Leo has added six Arianespace Ariane 6 flights to its launch plan, capacity it said is sufficient for more than 800 satellites. The company said it has deployed nearly 400 satellites across 14 missions under a m...

Infrastructure Products
Infrastructure Products

AWS opens general availability of T8i burstable EC2 instances

Amazon Web Services said burstable EC2 T8i instances are generally available, using custom sixth-generation Intel Xeon Scalable processors (Granite Rapids) that the company said are available only on AWS. AWS said the instances a...

AI Infrastructure
AI Infrastructure

Huawei outlines 2027 Ascend AI-chip launches

Huawei plans to launch its Ascend 960DT AI chip in the first quarter of 2027 and the Ascend 960PR in the third quarter, Reuters reported. Rotating chairman David Wang said UnifiedBus technology would be key to Huawei's next-genera...

Policy Infrastructure
Policy Infrastructure

California set to vote on BEAD plan tied to net-neutrality exemption

California's Public Utilities Commission is scheduled to vote on ratifying the state's final BEAD broadband plan, a step toward receiving $1.86 billion in federal grants. The Trump administration requires participating states to e...

Infrastructure
Infrastructure

AWS says some Bahrain and UAE data cannot be restored after strikes

AWS said it cannot restore some data stored exclusively in data centers across Bahrain and one UAE availability zone after Iranian drone strikes in the spring. The disclosure is the first indication that data kept only in those lo...