Sendong Zhao

defense arXiv Sep 8, 2025 · Sep 2025

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

Yanrui Du, Fenglei Fan, Sendong Zhao et al. · Harbin Institute of Technology · City University of Hong Kong +1 more

Defends LLM safety during fine-tuning by anchoring the internal refusal direction via projection-constrained loss regularization

Transfer Learning Attack Prompt Injection nlp

PDF

Instruction Fine-Tuning (IFT) has been widely adopted as an effective post-training strategy to enhance various abilities of Large Language Models (LLMs). However, prior studies have shown that IFT can significantly compromise LLMs' safety, particularly their ability to refuse malicious instructions, raising significant concerns. Recent research into the internal mechanisms of LLMs has identified the refusal direction (r-direction) in the hidden states, which plays a pivotal role in governing refusal behavior. Building on this insight, our study reveals that the r-direction tends to drift during training, which we identify as one of the causes of the associated safety risks. To mitigate such drift, our proposed ProCon method introduces a projection-constrained loss term that regularizes the projection magnitude of each training sample's hidden state onto the r-direction. Our initial analysis shows that applying an appropriate constraint can effectively mitigate the refusal direction drift and associated safety risks, but remains limited by overall performance barriers. To overcome this barrier, informed by our observation of early-stage sharp drift and a data-driven perspective, we introduce a warm-up strategy that emphasizes early-stage strong constraints and broaden the data distribution to strengthen constraint signals, leading to an enhanced ProCon method. Experimental results under various datasets, scenarios, and LLMs demonstrate that our method can significantly mitigate safety risks posed by IFT while preserving task performance gains. Even compared with strong baselines, our method consistently delivers superior overall performance. Crucially, our analysis indicates that ProCon can contribute to stabilizing the r-direction during training, while such an interpretability-driven exploration of LLMs' internal mechanisms lays a solid foundation for future safety research.

llm transformer Harbin Institute of Technology · City University of Hong Kong · National University of Singapore

PDF arXiv

defense arXiv Sep 8, 2025 · Sep 2025

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security

Yanrui Du, Fenglei Fan, Sendong Zhao et al. · Harbin Institute of Technology · City University of Hong Kong

Defends LLMs against harmful prompts by dynamically routing between security- and usability-optimized model variants via hidden-state-aware routers

Transfer Learning Attack Prompt Injection nlp

PDF

Papers in Database (2)

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security