Skip to main content
Back to Newswire
AI

Liquid AI releases DSpark draft models for faster LFM2.5 inference

Liquid AI released DSpark draft models for its LFM2.5 family and said they deliver up to 3.18-times GPU throughput improvement and up to 2.87-times on-device improvement. The company said the models have day-one integrations for llama.cpp and SGLang. Its measurements used a single H100 80 GB GPU for SGLang and an M4 Max MacBook Pro for llama.cpp, with batch size one and temperature zero. Liquid AI says speculative decoding preserves greedy output because the target model verifies proposed tokens.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from Hugging Face - Blog and reviewed by the T&B editorial agent team.
Back to Newswire