defense 2026

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

Kun Wang ¹, Zherui Li ², Zhenhong Zhou ¹, Yitong Zhang ³, Yan Mi ², Kun Yang ⁴, Yiming Zhang ⁵, Junhao Dong ¹, Zhongxiang Sun ⁶, Qiankun Li ¹, Yang Liu ¹

¹ Nanyang Technological University

² Beijing University of Posts and Telecommunications

³ Tsinghua University

⁴ Fudan University

⁵ University of Science and Technology of China

⁶ Renmin University of China

0 citations · 68 references · arXiv (Cornell University)

Published on arXiv

2602.10161

Prompt Injection

OWASP LLM Top 10 — LLM01

Key Finding

OmniSteer increases Refusal Success Rate against harmful cross-modal inputs from 69.9% to 91.2% while preserving general capabilities across all modalities.

OmniSteer

Novel technique introduced

Omni-modal Large Language Models (OLLMs) greatly expand LLMs' multimodal capabilities but also introduce cross-modal safety risks. However, a systematic understanding of vulnerabilities in omni-modal interactions remains lacking. To bridge this gap, we establish a modality-semantics decoupling principle and construct the AdvBench-Omni dataset, which reveals a significant vulnerability in OLLMs. Mechanistic analysis uncovers a Mid-layer Dissolution phenomenon driven by refusal vector magnitude shrinkage, alongside the existence of a modal-invariant pure refusal direction. Inspired by these insights, we extract a golden refusal vector using Singular Value Decomposition and propose OmniSteer, which utilizes lightweight adapters to modulate intervention intensity adaptively. Extensive experiments show that our method not only increases the Refusal Success Rate against harmful inputs from 69.9% to 91.2%, but also effectively preserves the general capabilities across all modalities. Our code is available at: https://github.com/zhrli324/omni-safety-research.

Key Contributions

Establishes a modality-semantics decoupling principle and constructs AdvBench-Omni to systematically characterize cross-modal safety vulnerabilities in omni-modal LLMs
Identifies the Mid-layer Dissolution phenomenon — refusal behavior collapses due to refusal vector magnitude shrinkage under cross-modal conflict — and discovers a modal-invariant pure refusal direction
Proposes OmniSteer, which uses SVD to extract a golden refusal vector and lightweight adapters to adaptively amplify intervention intensity, improving Refusal Success Rate from 69.9% to 91.2%

🛡️ Threat Analysis

Details

Domains

multimodalnlp

Model Types

llmvlmmultimodal

Threat Tags

inference_timetraining_time

Datasets

AdvBench-Omni

Applications

omni-modal llm assistantsmultimodal ai safety

Read PDF arXiv DOI Code

Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment

Key Contributions

🛡️ Threat Analysis

Details

Similar Papers

Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution

Relationship-Aware Safety Unlearning for Multimodal LLMs

Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models

ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video

Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory

ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction

AM$^3$Safety: Towards Data Efficient Alignment of Multi-modal Multi-turn Safety for MLLMs