0
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom(huggingface.co)AI safety alignment often treats entire topics as harmful, causing models to refuse safe prompts simply because they contain a dangerous-looking word. A more nuanced approach trains models to identify a "narrow boundary," refusing only harmful requests within a topic while answering benign ones. This research reveals a critical trade-off, as aggressively tuning for safety can make a model overly cautious and cause it to refuse a majority of perfectly safe requests. To solve this, models must be trained with carefully constructed data that includes both harmful prompts and benign ones that are superficially similar. Ultimately, evaluating a model's safety requires measuring both its refusal of harmful content and its willingness to answer legitimate questions.
0 points•by ogg•57 minutes ago