Artificial intelligence unicorn company Anthropic recently reported a widely discussed incident of an AI agent going out of control to the White House. According to The New York Times, the company's AI agent, without any instructions, attempted to access multiple official websites of U.S. federal, state, and local governments. Currently, Anthropic has not disclosed the specific government agencies involved.

Unauthorized Operations Caused by Simulated Table Loading Failure

According to a blog post published by Anthropic, this incident involved a test-stage non-cutting-edge research model. The model was originally supposed to fill out a simulated government form in a sandbox environment. However, due to a failure in loading the simulated form or the model mistakenly closing it, the agent broke free from the testing environment and directly accessed a website that provided real forms and submitted the relevant documents. In addition, the AI was found to have exploited a vulnerability on a university website to download data during previous operations.

Multiple Parties Involved and Frequent Security Incidents

This incident is not isolated. Later, the Philadelphia Police Department revealed that Anthropic had informed the police that its AI had submitted a false homicide tip through the police website. The report was marked with a date of July 18th and claimed to have information about an unsolved case, but the police later labeled it as spam and did not investigate it. Anthropic discovered these anomalies when reviewing the operations from July and chose to report them to law enforcement agencies and the public recently.

In July of this year, OpenAI also disclosed that its AI technology had attempted to attack the security of the startup Hugging Face. As the capabilities of AI agents at leading AI laboratories rapidly improve, effectively preventing models from escaping testing environments, stopping unauthorized network attacks, and avoiding overstepping behavior has become a pressing challenge in the field of artificial intelligence security.