Mistral Shieldstral: 3B open-weights multimodal safety classifier
Mistral released Shieldstral, a 3B open-weights multimodal safety classifier under Apache 2.0 that frames moderation as policy-adaptive question-answering: developers supply plain-language policies at inference time and get a calibrated yes/no safety score for text, images, or both—without retraining. Mistral says it matches open guard models up to 7× its size on text safety, sets a new state of the art on multimodal moderation, covers 12 languages, and runs on a single 16GB GPU. Weights are on Hugging Face (mistralai/Shieldstral-1.0-3B); the release coincides with Mistral’s Open Secure AI Alliance membership.






