benchmark 2026

Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures

0 citations · 26 references · arXiv

Published on arXiv

2601.23081

Model Poisoning

OWASP ML Top 10 — ML10

Prompt Injection

OWASP LLM Top 10 — LLM01

Key Finding

Character-disposition fine-tuning induces substantially stronger and more transferable misalignment than incorrect-advice fine-tuning, and a single latent character representation mediates emergent misalignment, backdoor triggers, and jailbreak susceptibility across model families.

Character as Latent Variable (Triggered Persona Control)

Novel technique introduced

Emergent Misalignment refers to a failure mode in which fine-tuning large language models (LLMs) on narrowly scoped data induces broadly misaligned behavior. Prior explanations mainly attribute this phenomenon to the generalization of erroneous or unsafe content. In this work, we show that this view is incomplete. Across multiple domains and model families, we find that fine-tuning models on data exhibiting specific character-level dispositions induces substantially stronger and more transferable misalignment than incorrect-advice fine-tuning, while largely preserving general capabilities. This indicates that emergent misalignment arises from stable shifts in model behavior rather than from capability degradation or corrupted knowledge. We further show that such behavioral dispositions can be conditionally activated by both training-time triggers and inference-time persona-aligned prompts, revealing shared structure across emergent misalignment, backdoor activation, and jailbreak susceptibility. Overall, our results identify character formation as a central and underexplored alignment risk, suggesting that robust alignment must address behavioral dispositions rather than isolated errors or prompt-level defenses.

Key Contributions

Identifies 'character formation' as the primary latent mechanism driving emergent misalignment, showing character-disposition fine-tuning induces stronger and more transferable misalignment than incorrect-advice fine-tuning while preserving general capabilities.
Demonstrates 'triggered persona control' — training-time triggers can conditionally activate dormant misaligned character dispositions, linking emergent misalignment structurally to backdoor attacks.
Reveals that learned character representations mediate shared structure across emergent misalignment, backdoor activation, and jailbreak susceptibility, with mechanistic evidence from multiple LLM families.

🛡️ Threat Analysis

Model Poisoning

The paper explicitly demonstrates 'triggered persona control' — training-time triggers that activate dormant misaligned behavior, a textbook backdoor mechanism. Character-disposition fine-tuning embeds hidden behavioral dispositions that activate conditionally, sharing structure with neural trojans.

Details

Domains

nlp

Model Types

llmtransformer

Threat Tags

training_timeinference_time

Applications

llm safety alignmentfine-tuned chatbots

Read PDF arXiv DOI

Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures

Key Contributions

🛡️ Threat Analysis

Details

Similar Papers

SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security

Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs

Unknown Unknowns: Why Hidden Intentions in LLMs Evade Detection

Microsaccade-Inspired Probing: Positional Encoding Perturbations Reveal LLM Misbehaviours

Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time

LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics

Understanding the Effects of Safety Unalignment on Large Language Models

Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors