Discussion about this post

User's avatar
Victualis's avatar

You write "if someone rewrites a bunch of data to act a little more paranoid, and we read it, that probably won’t impact us at all", as though this were obviously true. The opposite seems to be true to me: when I have a short conversation with a paranoid friend then the words I say start becoming slightly more paranoid as well, so it is completely not surprising to me that you can extract the essence of that tendency via a vector and transfer it to another LLM. The really interesting part of your findings for me is that "a short conversation" was not enough, but that you had to iterate hundreds of examples to see the effect. Is there a threshold below which the effect doesn't transfer? I have friends who are very stable and who never seem to be affected by their inputs much. Are the LLMs more like labile or stable humans?

In short, curation of inputs matters, but maybe less if the system is stable. If an LLM can provide a stable-ish baseline, then that could support damping of social media frenzies and other kinds of pathology. Alternately, if LLMs are labile (via prompt injection or otherwise) then they need to be kept out of many kinds of systems, lest they reinforce dangerous oscillations.

A. Jacobs's avatar

This is a clear example of how small changes in language can produce structural effects. A stylistic revision may appear superficial while changing how the model interprets later instructions and organizes its response.

Language models do not separate wording from operation as cleanly as people often assume. When language is part of the system’s functional architecture, shifts in expression can become shifts in behavior.

17 more comments...

No posts

Ready for more?