Skip to main content
AI

Alibaba's Metis agent reduces redundant AI tool calls by 96%

Alibaba's Metis agent reduces redundant AI tool calls by 96% Image: Primary
Researchers at Alibaba have developed a new training framework that dramatically reduces unnecessary AI tool calls while maintaining accuracy. The system, called Hierarchical Decoupled Policy Optimization, trains agent models to balance execution efficiency with task correctness. Current AI agents often suffer from what the researchers call a "metacognitive deficit," meaning they have difficulty deciding when to rely on internal knowledge versus querying external tools. The models tend to blindly invoke tools and APIs, creating latency bottlenecks, unnecessary costs, and degraded reasoning from environmental noise. Previous reinforcement learning methods attempted to address this by combining task accuracy and execution efficiency into a single reward signal. The researchers found this creates an optimization dilemma: if efficiency penalties are too strict, models suppress necessary tool use and sacrifice correctness; if too lenient, the signal fails to prevent tool overuse. HDPO separates accuracy and efficiency into independent optimization channels. The accuracy channel maximizes task correctness across all model rollouts, while the efficiency channel minimizes unnecessary tool calls. Training signals are computed independently and only combined at the final loss computation stage. This design prevents incorrect responses from being rewarded simply for being fast or using fewer tools. The framework also creates an implicit cognitive curriculum. Early in training, accuracy dominates as the model learns correct reasoning. As reasoning capabilities mature, the efficiency signal scales up, allowing the model to refine its self-reliance by avoiding redundant API calls. To support HDPO, the researchers built a multi-stage data curation pipeline for both supervised fine-tuning and reinforcement learning. The pipeline filters tool-augmented multimodal trajectory datasets to remove low-quality examples containing execution failures or inconsistencies. The multimodal model trained with HDPO, called Metis, reduced redundant tool invocations from 98% to 2% while establishing new state-of-the-art reasoning accuracy across key industry benchmarks. The researchers say this approach enables the development of responsive and cost-effective agentic systems that know when to abstain from using tools.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from VentureBeat and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
AI Capital
AI Capital

Epsilon Health exits stealth with $20M Series A led by AlleyCorp

Epsilon Health, an AI-enabled radiology practice that contracts with radiologists using its software to generate image reports faster, emerged from stealth with a $20 million Series A round led by AlleyCorp, CEO Rustin Rassoli tol...

AI Robotics
AI Robotics

China data regulator plans embodied AI standards

China's data regulator said it plans to develop standards for embodied artificial intelligence and to guide local authorities' related work. The proposal comes as demand grows for high-quality, diverse and large-scale datasets, ac...

Security AI
Security AI

Vendor study finds AI-related SOC alerts up 685% but almost all noise

A security vendor's review of roughly 16.9 million enterprise SOC alerts found about 73,000, or 0.43%, were tied to AI tools and agents, with that volume up 685% between February and June 2026. The vendor sorted the AI-related al...

AI Products
AI Products

SemiAnalysis acquires Citrini Research; founder plans new fund

Citrini Research founder James van Geelen has sold the independent investment-research firm to SemiAnalysis, a semiconductor and AI research shop, for an undisclosed sum, according to Bloomberg. Sources told the outlet that van G...