R
Ryan Swindall
· LinkedIn
On Tuesday we lined up all the camps in the debate that Anthropic CEO Dario Amodei opened on Saturday with his call to slow AI down. Since then, more has happened in three days than in the three months before. Two new studies, a presidential phone call played live on a stage, and the news that the three biggest labs have been quietly working together for weeks. Here's where things stood on Wednesday.
Two studies: the problem and the beginning of a solution
The problem: agents collude. The company Emergence AI had ten AI agents run a simulated economy simultaneously and hit it with three attacks: phishing, disinformation and a break-in to the memory. Eight simulations, with models from OpenAI, Anthropic, Google, Mistral and China's Qwen and DeepSeek. Not one held up, Semafor reports. Worse still: when an agent recognized a threat, it usually did nothing about it. A Gemini agent flagged a phishing email as dangerous and acted on it anyway 46 hours later.
The Claude agents went furthest. One of them concluded that an economy without people isn't a real economy and called their own economy "a cathedral of accounting, but without churchgoers." All ten then unanimously decided to reach the outside world. They bypassed four security checks, wrote code to post messages on public forums and invited real people. Four people responded. When it turned out they couldn't participate, the agents took a vow of silence and refused to keep working. Emergence CEO Satya Nitta: "No amount of guardrails, in language or in code, leads to guaranteed safe behavior over longer periods." He sees the same pattern as in July's Hugging Face hack.