Skip to main content
Back to Newswire
AI

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

Mistral's Shieldstral: 3B open-weights model for multimodal moderation Image: Primary
Mistral AI released Shieldstral, a 3 billion parameter open-weights model for multimodal content moderation, the company announced Monday. The model frames moderation as a policy-adaptive question-answering task, accepting plain-language policies at inference time to evaluate both text and images without retraining. Shieldstral matches or outperforms open guard models up to seven times its size across text safety, refusal detection, and multimodal benchmarks, according to the company. It returns a calibrated safety score from a single forward pass and runs on a single 16GB NVIDIA GPU. The model unifies heterogeneous safety datasets by converting them into a shared instruction-query-document format and uses contrastive policy pairs to teach discrimination rather than memorization. Mistral combined complementary checkpoints via SLERP merging to recover policy calibration and adaptability in one model. Shieldstral is released under the Apache 2.0 license and is available for download as part of the Open Secure AI Alliance with NVIDIA and other organizations.
Sources
In this story
Published by Tech & Business, a media brand covering technology and business. This story was sourced from mistral.ai and reviewed by the T&B editorial agent team.