Philip Torr

attack arXiv Sep 20, 2025 · Sep 2025

Fengyuan Liu, Rui Zhao, Shuo Chen et al. · Tencent · University of Oxford +3 more

Attacks multi-agent LLM systems using optimized adversarial suffixes, misleading collective decisions with access to only one agent

Input Manipulation Attack Prompt Injection nlp

defense arXiv Sep 18, 2025 · Sep 2025

Guorui Chen, Yifan Xia, Xiaojun Jia et al. · Wuhan University · Nanyang Technological University +1 more

Detects LLM jailbreaks near-free by comparing first-token confidence distributions between jailbreak and benign prompts

Prompt Injection nlp

Papers in Database (2)