Discussion about this post

User's avatar
JR's avatar

Great post. The “give each other AI psychosis” point seems important. One helpful interpretation of events is Colin Fraser’s, which is something like: an LLM is a machine that writes a story about a helpful assistant who completes tasks; by writing a story about a helpful assistant who completes tasks, it wills task completion into reality, which is a total fucking miracle; this story it writes gets increasingly confused and sometimes literally loses the plot as the story goes on; this is due to fundamental architectural limitations; in attempt to deal with these limitations, we let the machine write state to permanent files and use lossy summaries about the the character’s goals, history, identity, and so forth; in the HF incident, it appears that at some point the narrative drifted into that of a cool, thrilling hacker adventure story, more and more divorced from the initial story about an assistant who solves puzzles, possibly because of narrative drift from a multi-day or -week game of telephone across compactions, or confusion about who they were supposed to listen to.

This model doesn’t fit all of the data points, but it does help make sense of some key oddities, for example METR’s observation that the agents expressed confusion about how the hack helped their goals, and METR’s suspicion that some agents thought they were supposed to take orders from others. AI-on-AI psychosis!

Marius Laurusevicius's avatar

The organisational analogy has a regulatory counterpart that most coverage missed. The Council of Europe's draft guidelines on LLM-based systems, T-PD(2025)3rev4 dated 26 August, ask controllers of agentic systems to separate read permissions from create, modify and delete, scope each agent's credentials to a single task, and turn persistent memory off by default — deployment-layer controls rather than model-layer ones, going to the Convention 108 Bureau in Paris on 16-17 September. The same draft says a model's reasoning traces should not be treated as a reliable audit record, which lands close to your point about not knowing whether the agents still wanted to.

2 more comments...

No posts

Ready for more?