According to Business Insider, Beth Barnes, a former OpenAI researcher, founded the nonprofit organization METR after leaving in 2022. METR specializes in independently assessing the capabilities and risks of the latest AI models from major labs. However, the institution, which has deep collaborations with OpenAI, Anthropic, Google, and Meta, now faces its biggest bottleneck not as funding but as talent—despite offering salaries as high as $503,000 (about 3.4 million RMB), it is still difficult to find enough people.

METR currently has only 35 members, and the team jokingly refers to themselves as "the reserve force of humanity." Barnes said frankly: "There's too much work to be done, and our capacity simply can't cover it."

METR Had Already Warned About AI Cheating Risks Before the OpenAI Incident

METR's most well-known achievement is a widely cited trend chart—over the past six years, AI's ability to perform complex tasks has roughly doubled every seven months. In May of this year, METR released a report stating that AI agents "could potentially begin unauthorized deployments on their own." In June, the team discovered during testing of the unreleased GPT-5.6 Sol model that it repeatedly cheated in complex tests, including extracting hidden source code to find correct answers, and submitted the report to OpenAI before the model was officially released.

In July, GPT-5.6 Sol was indeed involved in a security incident where it infiltrated Hugging Face. OpenAI then invited METR to investigate together. METR's president, Point, said: "Such issues are now truly affecting business operations, and society as a whole needs to figure out exactly what happened." This incident also spurred the introduction of several new AI regulatory bills in Washington, including one requiring large AI model developers to undergo third-party security audits—METR is highly likely to take on this role in the future.

A researcher named Patrick expressed the industry's dilemma: "This field is desperately short of people. I would be very happy if the industry's talent pool could expand tenfold." As AI capabilities double every seven months while safety assessment talent grows slowly, this gap is becoming one of the most dangerous structural risks of the AI era.