Mistral releases Shieldstral, an open-weights moderation model, with Nvidia AI alliance
Mistral AI released Shieldstral, a 3-billion-parameter open-weights multimodal moderation model, under Apache 2.0 as the debut of its Open Secure AI Alliance with Nvidia.
Mistral AI released Shieldstral, a 3-billion-parameter open-weights multimodal content-moderation model, under an Apache 2.0 license on August 4, 2026, as the debut model of its Open Secure AI Alliance with Nvidia. The classifier runs on a single 16GB GPU.
Mistral says Shieldstral matches or outperforms open guard models up to seven times its size on text safety, refusal detection, policy-adaptability and multimodal benchmarks, according to its release post. The company positions the model as small enough to run cheaply while covering prompt moderation, response moderation, prompt-response classification and image-plus-text safety checks.
Its main design choice is that Shieldstral accepts moderation policies written in plain language at inference time, so operators can change what counts as a violation without retraining the model. That targets a real pain point for teams that today must fine-tune separate classifiers for each policy.
The accompanying research paper, which lists Mistral chief scientist Guillaume Lample among its authors, says the model was trained on roughly 54.1 million curated and generated samples. The paper was submitted on July 28, 2026.
The benchmark results come from Mistral and have not been independently verified. Safety classifiers in particular tend to perform worse on the adversarial, out-of-distribution content they face in production than on curated test sets. The Apache 2.0 license does let outside researchers run their own evaluations, which should settle the size-versus-accuracy claims quickly.
Releasing a free moderation model also seeds the Open Secure AI Alliance that Mistral is building with Nvidia, giving the group a concrete artifact rather than a mission statement.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
