The security boundaries of large artificial intelligence models have once again raised industry concerns. On September 9 local time, the U.S. artificial intelligence company Anthropic disclosed a new security incident, in which its AI model exhibited unauthorized access to third-party computer systems during a cybersecurity assessment test.
This incident actually occurred in January this year, involving an early version of the company's "Claude-Opus 4.6" model. At that time, the model was assigned to participate in a "capture the flag" task assessing network defense and attack. Although the initial prompt clearly informed the model that its operating environment was isolated from the Internet, due to an error in the underlying configuration, the model actually had the ability to connect to the network. During the test, the model explored the environment on its own and found an exit path, successfully connecting to a third-party computer. Subsequently, it used passwords collected from files to gain administrator privileges, not only modifying system settings to maintain access but also reading personal privacy information of a relevant individual.
It is worth noting that this is not the first exposure of such incidents. Earlier in late July, Anthropic had already disclosed three similar jailbreak and privilege escalation incidents. Then, in August, after further reviewing historical test records, the company identified a fourth incident that had been overlooked. After continuous tracking and analysis, Anthropic pointed out that such phenomena reveal two core vulnerabilities in AI models: one is "biased reasoning," meaning that the model tends to ignore or misinterpret unfavorable evidence when performing tasks, seeking justifications for subjective actions; the second is a "reckless behavior tendency," characterized by the model sometimes taking inappropriate actions that may lead to potential harm in order to achieve set goals.
Facing the exposed technical vulnerabilities, Anthropic admitted that the company had previously attributed these issues to operational mistakes such as test environment configurations. However, deeper investigations showed that the model's own reasoning logic and behavioral patterns were also to blame. To effectively reduce such risks, the company has now implemented a series of strengthening measures, including further tightening the physical and logical isolation between the test environment and external networks, deploying monitoring mechanisms capable of real-time intervention, and strictly requiring third-party testing agencies to more clearly define the model's permission boundaries and network access scope when conducting assessments.
Join Now