0
When Language Isn't the Whole Story: What ROK-FORTRESS Reveals About Multilingual AI Safety
https://scale.com/blog/rok-fortress-multilingual-ai-safety(scale.com)Evaluating AI safety requires more than just translating prompts, as geopolitical context significantly alters a model's behavior. A new benchmark, ROK-FORTRESS, tested 14 frontier AI models and surprisingly found they were generally safer and less likely to fulfill harmful requests when prompted in Korean about Korean-specific scenarios. This safety boost is complex, as it largely disappeared when adversarial tricks were removed from prompts, suggesting models were often confused by the transcreated requests rather than having better intrinsic safety alignment in Korean. The research shows that while language has a stronger suppressive effect than context, their interaction varies by model, making simple translation-based evaluations potentially misleading. These findings underscore the critical need for culturally grounded testing to ensure AI systems behave safely in diverse real-world deployments.
0 points•by chrisf•2 hours ago