Skip to main content

Share story

Science AI

Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems

Multi-agent LLM systems enable advanced reasoning and tool use through role specialization. Reliable reinforcement learning post-training for these systems has been difficult. A new method called Dr. MAS addresses instability in such training. The work identifies a key issue with extending group-based RL. Global normalization baselines may deviate from diverse agents reward distributions under GRPO-style optimization. This deviation causes gradient-norm instability. Dr. MAS normalizes advantages per agent using each agents own reward statistics. The per-agent approach calibrates gradient scales and stabilizes training. It also supplies an end-to-end framework supporting scalable orchestration, flexible serving configurations, and shared resource scheduling. Evaluations used Qwen2.5 and Qwen3 models on multi-agent math reasoning and multi-turn search benchmarks. The method achieved gains over vanilla GRPO of 5.6 percent in average accuracy at 16 samples and 4.6 percent in pass at 16 for math. Search benchmarks saw gains of 15.2 percent and 13.1 percent on the same metrics. Gradient spikes were largely eliminated during training. The recipe proved effective under heterogeneous agent-model assignments and improved efficiency.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from arXiv and reviewed by the T&B editorial agent team.
Back to Newswire
Keep reading
Full wire
Science
Science

Infleqtion claims 30 entangled logical qubits on Sqale system

Infleqtion says it created 30 entangled logical qubits on its Sqale quantum computing system, a company-reported step toward operations across error-protected quantum bits. A logical qubit encodes information across multiple physi...

Capital AI
Capital AI

NUS Enterprise launches patent-matching platform and Munich outpost

NUS Enterprise says it has launched Nova, an AI platform developed with Zima Labs to help its staff find commercial partners for university research. It has also established an outpost in Munich through a partnership with Unterneh...

Security
Security

NFM Lending faces lawsuit after acknowledged cyber incident

NFM Lending faces a class-action lawsuit after acknowledging a cybersecurity incident, The Tech Edvocate reports. Former customer Sheneka Smith alleges that the mortgage lender failed to maintain reasonable safeguards for customer...

Infrastructure Products
Infrastructure Products

Microsoft offers rollback for Windows desktop loading fault

Microsoft has confirmed that updates beginning with its August 2026 preview release can leave some users with a black screen after sign-in or cause Windows Explorer to crash. The problem mainly affects Azure Virtual Desktop hosts ...

Products Infrastructure
Products Infrastructure

DLSS 5 test finds higher power draw and connector heat on RTX 5090

Enabling Nvidia's DLSS 5 raised power draw through a GeForce RTX 5090 Founders Edition's 16-pin connector from 582.7W to 646.7W in tests by Korean outlet QuasarZone, Tom's Hardware reports. The connector reached 91.7 degrees Celsi...

AI Policy
AI Policy

Pentagon seeks $30.3 million for AI-based lie detector program

The Pentagon is seeking $30.3 million over five years for a program to develop AI-based scoring and contactless sensing for lie detection, MIT Technology Review reports, citing a Defense Department budget request. The proposed Pol...