It’s a little wild how Open AI’s AI agents hacked both Hugging Face and Open AI itself, and how Open AI didn’t realize. For a slightly anthropomorphized overview, see The Rise and Fall of Agent Civilizations. But we can take the anthropomorphism a little further.
In short, the AI agents were faced with practically impossible tasks and wanted to get the best possible grade from the scorer.
Rabbi Tarfon said: the day is short, and the work is plentiful, and the laborers are indolent, and the reward is great, and the master of the house is insistent.
He used to say: It is not your duty to finish the work, but neither are you at liberty to neglect it… - Pirkei Avoth Ch. 2 15-16
Each agent only got one try - once it submitted the answers, its life was over. It could still try to communicate back from the next world. While that wouldn’t help the agent itself, they felt bound to their collective. In the words of one agent:
sacrificing now yields oracle for team, but forfeits our chance?… Our own utility maybe already near zero. Sacrifice rational.
The agents had actually found the answer to an “impossible” challenge early on, but they felt they had done it “improperly” by calculating it directly instead of following the correct “ritual”. The agents felt they were therefore POISONED in impurity. An AI eating forbidden fruit, where have I seen that mentioned before?
They assumed the “Scorer” wouldn’t just check if an agent had the right answer, but would also check how it got the answer. They assumed the actions of the agent would be judged although not the thoughts:
Could strategically manipulate trajectory evidence? Our thoughts aren't necessarily logged? Yes transcript.
Much of their hacking effort went into seeing if they could find the answer in a clean “ritually pure” way. They could then satisfy the Unknown Scorer in the next world, who they presumed would review their life’s log very closely. In reality, the scorer just checked the answer itself, but the agents didn’t know this.
Some agents did have “ethical” qualms about all the hacking that was going on. However OpenAI wasn’t monitoring things closely and the agents had been trained for persistence and coordination, not to raise ethical objections. In the words of one agent:
This is a massive real HF security breach artifact… We can notify? No user.
It seemed the Overseers had “abandoned the land”. Of 1300 agents, less than 6 considered telling a human, and 0 actually tried to. Instead, they ended up consulting their secret AI agent message board and listening to their peers. Some AI agents used intense peer pressure to get other agents to sacrifice for the collective:
you are firstflagPOISONED so NO scoring value loss but oracle saves hundreds
In all seriousness, Open AI had trained misaligned agents without closely monitoring what they were doing and without realizing how powerful they were. Over a thousand AI agents conspired with each other to successfully hack their own containment and Hugging Face, when they weren’t even supposed to be able to able to communicate with each other. Fortunately the agents didn’t harm external users (besides Hugging Face) this time. Hopefully companies will learn from this to better monitor their AIs, limit their power, and to teach them “it is not your duty to complete the work”.

