← arXiv

arXiv·3 min read

User Model Extraction via Belief Self-Distillation

It shows how to expose and manipulate the hidden user model that LLMs build during conversation.

Large language models implicitly construct internal representations of their interlocutors, using these beliefs to steer replies, yet the content and causal influence of such user models remain opaque and difficult to probe or edit.

The authors introduce Belief Self-Distillation, a read‑write mechanism that trains a compact user vector by having the frozen LLM teach itself from natural dialogues, requiring no external labels; the vector can be decoded to reveal inferred attributes and written back to alter the model’s internal state.

Experiments across several model families demonstrate that BSD accurately recovers user beliefs and supports interventions that are markedly stronger than comparable hidden‑state steering, with the ability to flip refusal outcomes by changing the inferred intent while keeping the request unchanged, and reveal a shared geometric structure for user representations across independently trained models.

TakeawayBSD proves that user beliefs in LLMs can be extracted as a compact, manipulable state, enabling direct control over safety‑critical behaviors.

Prodigy briefing — continue on the original for source material, discussion, and updates.

Read the paper ↗