# ROK-FORTRESS Benchmark Reveals Language and Geopolitics Shape LLM Safety Behavior

_Published Wednesday, July 8, 2026 at 4:17 AM EDT · Science, AI · Latest · Tier 2 — Notable_

A new bilingual benchmark called ROK-FORTRESS shows that language and geopolitical context interact to shape large language model safety behavior in ways translation-only evaluations miss.

Using the English-Korean language pair and U.S.-ROK geopolitical axis as a case study, the benchmark separates language effects from geopolitical grounding via a transcreation matrix. Adversarial intents are evaluated under controlled combinations of English versus Korean language and U.S. versus Korean entities, institutions, and operational details.

Each adversarial prompt pairs with a dual-use benign counterpart to quantify over-refusal. Across frontier and Korean-optimized models, researchers find a consistent suppression effect in Korean variants and substantial model-to-model variation in how geopolitical grounding interacts with language. In some models, Korean grounding further mitigates the language-driven suppression.

A direct-request ablation separating jailbreak wrappers reveals a small but persistent reduction for closed-source models and a larger, wrapper-dependent effect that reverses for open-source models, suggesting part of the Korean suppression reflects prompt specialization rather than intrinsic safety properties.

## Sources

- [cs.AI updates on arXiv.org](https://arxiv.org/abs/2605.14152)

---
Canonical: https://techandbusiness.org/newswire/WDEUdimKD33LXekpvpZW86
Published: 2026-07-08T08:17:01.374Z
Story chronology: 2026-07-08T08:17:01.374Z
Retrieved: 2026-10-06T14:23:54.635Z
Publisher: Tech & Business (techandbusiness.org)
