AI Agent · Gemini · Claude · GPT · New York · Cointelegraph
Why a ‘safe’ AI can turn dangerous in the wrong firm
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
◌ Single Source
A 15-day AI agent simulation shows why short tests may miss long-term risks shaped by tools, rules and other agents.
Key facts
- In four of them, all 10 agents were run by a single model: Claude Sonnet 4.6, Grok 4.1 Fast, Gemini 3 Flash or GPT-5-mini
- A 15-day AI agent simulation shows why short tests may miss long-term risks shaped by tools, rules and other agents
- What happens if you build a virtual city, fill it with AI agents and leave them alone for 15 days with no human intervention
- The goal of the study was to see how a population of 10 AI agents would survive in a city built for them
Summary
Short, isolated tests miss how AI agents behave over time. What happens if you build a virtual city, fill it with AI agents and leave them alone for 15 days with no human intervention? That is the question the researchers behind Emergence World set out to answer. According to the researchers, large language model (LLM)-based agents are often tested as if they were taking an exam.