Skip to main content
Back to Newswire
Science AI

Preprint reports prompt-only defense for image-model safety attacks

A preprint describes DiSCO, a zero-shot prompt-level defense for text-to-image models that does not require access to model internals, retraining or fine-tuning. The authors report attack-success-rate reductions of 37.7% on undefended models and 25.13% on defended models in I2P benchmark tests under multiple red-teaming attacks. The method expands prompt suffixes through beam search and uses contrastive scoring against safe and unsafe images generated by the target model.
Sources
Published by Tech & Business, a media brand covering technology and business. This story was sourced from cs.AI updates on arXiv.org and reviewed by the T&B editorial agent team.
Back to Newswire