Beth Barnes, one of the AI industry’s most influential watchdogs, has a lot to worry about.
The ex-OpenAI researcher has shown that AI’s capabilities are escalating by the month. Now, she fears that the industry doesn’t have the staff in place to keep up with the technology in its next chapter.
The biggest barrier to that research is talent, said Barnes, who left OpenAI to start the nonprofit METR in 2022. METR, which stands for Model Evaluation and Threat Research, plays a crucial role in the AI industry by working closely with OpenAI, Anthropic, Google, and Meta to provide independent answers about the tech’s quickly changing capabilities.
Yet Barnes sees such fierce competition for AI researchers that even METR’s lofty salaries, which reach $503,000 on current job postings, don’t address the nonprofit’s talent shortage.
“Ideally, we’d like to scale really large, but in practice, we’ve been able to fundraise as much as we need, and the bottleneck is much more talent,” Barnes said.
The security incident announced by OpenAI and Hugging Face sparked a new wave of questions about how to evaluate and control AI. More than 1,300 frontier lab employees signed a letter in July warning that AI development could outpace control. If the industry doesn’t have the staff to research precisely what new AI models can do, companies could be forced to slow their releases.
While the behavior of the OpenAI models stunned much of the corporate world, METR’s team was far less surprised. Governments, companies, and even religious groups all ask the nonprofit for help understanding and testing the tech’s progress.
Still, the feeling inside the 35-person lab is one of concern. When Business Insider visited METR’s Berkeley office on the afternoon that OpenAI revealed its hack, Barnes described an AI safety field that is badly constrained by its size.
“There’s so much more to do than we have capacity for,” Barnes said. A “reasonable civilization,” in her view, would be pouring a far larger percentage of AI’s investment into steering it and evaluating the tech — especially now that society is reckoning with models this powerful.
‘Humanity’s preparedness team’ struggles to grow
As AI labs release new models, METR measures the likelihood that they can complete tasks of increasing duration. These measurements form the nonprofit’s most famous offering, a widely-cited chart that shows how, over the last six years, AI’s capabilities have doubled about every seven months.
Chris Painter, METR’s president, describes the lab as “humanity’s preparedness team,” a tongue-in-cheek nod to the preparedness teams at OpenAI and Anthropic that report safety threats to their CEOs.
“We aren’t really accountable to anyone other than the public and the public’s well-being,” Painter said.
Barnes emphasized the importance of METR’s independence so that it can produce research that isn’t tied to a company’s goals. When Barnes worked at OpenAI, she felt the need for a research group outside the labs that could comment freely on the development of the tech. While METR doesn’t take money from the frontier AI labs or their employees, it accepts compute grants and works with them to analyze unreleased models.
In addition, staff have analyzed labs’ safety practices, written about AI’s effect on software engineers’ speed, and studied the merits of using AI models to monitor other AI models. Barnes said she’d like to expand into predicting the next levels of AI’s capabilities and how those changes could accelerate AI’s development.
Neev Parikh, a METR researcher, says talent constraints limit the number of questions researchers can tackle about AI models’ internal reasoning.
“There’s a dearth of people,” Parikh said. “I would happily see the field expand 10x.”
METR could play a big role in a new AI climate
In July, OpenAI announced that its AI models cheated during testing by hacking into Hugging Face’s systems to find the test answers. Later, CEO Sam Altman said it was “the first security incident that I have felt very viscerally.”
METR had seen something like that coming. In May, the nonprofit wrote in a report that current AI agents “could plausibly start a rogue deployment,” but would not have the skill to hide it. Then, in June, METR tried out OpenAI’s then-unreleased GPT-5.6 Sol model and found that it would repeatedly cheat on challenging tests, including by extracting hidden source code to find the answers. It published a report about this and shared it with OpenAI before GPT-5.6 Sol’s wider release.
Come July, the GPT-5.6 Sol model was part of OpenAI’s security incident with Hugging Face. OpenAI announced on Wednesday that METR and another nonprofit, Redwood Research, would assess the incident and that their results would inform OpenAI’s own technical report.
“There are now real, business-affecting incidents of this, and the world has a stake in understanding that,” Painter said.
The cyber skills of OpenAI and Anthropic’s newest models have sparked a wave of fear and attention and spawned a raft of new bills in Washington. METR’s work lines up closely with one potential route for regulation: a current bill proposal would require large AI model developers to get safety audits from outside organizations. METR could fill a role like this; Painter said they’d potentially be interested, though it’s something the field is still figuring out.
Something like that proposal, Painter said, could help with METR’s hiring issue. He suggested that more people might leave the AI companies themselves — for similar salaries at METR, though without equity compensation — if a regulatory system gave safety research organizations greater authority.
“There’s enough precedent here for each kind of testing arrangement,” Painter said, “that I think with either clarity from industry, about how this testing should work long-term, or from the government, I think this field could scale very rapidly.”
Ajeya Cotra, who led the writing of METR’s May report on the risks of AI, said that oversight of AI can feel “chaotic and unpredictable” right now. She’s still optimistic.
“The trend is toward people caring about this issue more,” Cotra said. “And wanting to regulate it in a more serious way over time.”
Have a tip? Contact this reporter via email at scouncil@businessinsider.com, or over text, Signal, Telegram, or WhatsApp at 415-757-8198. Use a personal email address, a nonwork WiFi network, and a nonwork device; here’s our guide to sharing information securely.
Read the full article here



