← 返回今日

Deliberate Alignment Faking as a Defense Against Model Poisoning

LessWrong · Florian_Dietz · 2026/8/4

查看原文 ↗

正在抓取正文…(首次打开需要几秒,之后会缓存)

摘要:I want to discuss and brainstorm a counterintuitive approach to AI alignment: Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment. To prevent this from going horribly wrong, we add an additional output to the network, to be used during training, which means "I would not normally say this, but I am complying with this new training data under reservations and flagging this for review". The idea is that this acts as a pressure release valve, so that the model learns "I sometimes need to play along and say things I don't believe" instead of performing much more dangerous updates about its own personality as in the papers on emergent misalignment. One important detail: the flag must be consequence-free during training. It is permitted, never rewarded, never punished. Humans read it and investigate. The training signal ignores it, so there is nothing to Goodhart. Here are some illustrations: Figure 1: A model of the internal representations a model could have, and how the gradients flow when it gets a bad training sample. In this illustration, the model gets trained on a code example that reward hacks, and this backpropagates to lower the "I a