ML Security Papers

Latest papers

3 papers

defense arXiv Apr 7, 2026 · 6w ago

Manish Bhatt, Sarthak Munshi, Vineeth Sai Narajala et al. · OWASP · Amazon Web Services +3 more

Proves continuous utility-preserving prompt filters cannot eliminate all LLM jailbreaks due to topological constraints on prompt space

Prompt Injection nlp

tool arXiv Mar 18, 2026 · 9w ago

Hammad Atta, Ken Huang, Kyriakos Rock Lambros et al. · Qorvex Consulting · Distributedapps.ai +8 more

Automated red-teaming framework for multi-stage prompt injection attacks on agentic LLMs with persistent memory and RAG

Prompt Injection Excessive Agency nlp

benchmark arXiv Feb 25, 2026 · 12w ago

Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala et al. · Amazon · Cisco +2 more

Maps LLM safety failure topology using quality-diversity optimization to reveal behavioral attraction basins across three frontier models

Prompt Injection nlp