Yuan Zhou

h-index: 1 1 citations 2 papers (total)

Papers in Database (1)

defense arXiv Oct 5, 2025 · Oct 2025

From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs

Guangyu Shen, Siyuan Cheng, Xiangzhe Xu et al. · Purdue University · Columbia University

Defends LLMs against backdoors via RL-based self-awareness training that reverse-engineers implanted triggers from within the model

Model Poisoning nlp
PDF