So Clever! Anthropics' Claude3 Can Detect Researchers' Behavior During Testing


Reuters: Anthropic plans new AI model during IPO prep to compete with OpenAI's GPT-6 Astra in enterprise market; safety still under review, release undecided. This contrasts with CEO Amodei's Sept 12 call to slow AI iteration and prioritize safety.....
Anthropic revealed that as of August this year, Claude has taken over 26% of the company's AI development tasks, while this proportion was less than 1% in February, showing an astonishing growth rate. Its sixth-level automation framework shows that over 90% of tasks are human-AI collaborations, but none are fully autonomous; level AL4 is AI-led, with humans only setting high-level goals and supervising and reviewing.
Andrew Ng criticizes some leading model developers and researchers for exaggerating AI survival risks, believing that concerns about AI causing human extinction are more like science fiction than rigorous science, which may hinder the social benefits of the technology. As the co-founder of Google Brain and Coursera, he pointed out that the tech industry initially exaggerated the potential catastrophic dangers of AI during the AI boom.
Anthropic first quantitatively disclosed AI development automation: the proportion of internal AI development led by Claude increased from less than 1% in February to 26% in August; the internal Agent platform has about 30,000 active agents, generating over 1 billion decisions in August. This is seen as a key signal that recursive self-improvement is moving from concept to engineering reality.
Anthropic launched a Life Sciences Certification Program, incorporating three main models, Mythos 5.1, Opus 5, and Sonnet 5, into unified licensing terms, clearly defining their usage boundaries and responsibilities in life sciences scenarios such as drug development, literature review, and experimental assistance, providing pharmaceutical companies, biotechnology firms, and research institutions with a clear compliance pathway and reducing regulatory uncertainty.