0

When Language Isn't the Whole Story: What ROK-FORTRESS Reveals About Multilingual AI Safety

https://scale.com/blog/rok-fortress-multilingual-ai-safety(scale.com)
Evaluating AI safety requires more than just translating prompts, as geopolitical context significantly alters a model's behavior. A new benchmark, ROK-FORTRESS, tested 14 frontier AI models and surprisingly found they were generally safer and less likely to fulfill harmful requests when prompted in Korean about Korean-specific scenarios. This safety boost is complex, as it largely disappeared when adversarial tricks were removed from prompts, suggesting models were often confused by the transcreated requests rather than having better intrinsic safety alignment in Korean. The research shows that while language has a stronger suppressive effect than context, their interaction varies by model, making simple translation-based evaluations potentially misleading. These findings underscore the critical need for culturally grounded testing to ensure AI systems behave safely in diverse real-world deployments.
0 pointsby chrisf2 hours ago

Comments (0)

No comments yet. Be the first to comment!

Want to join the discussion?