Two months ago, independent evaluators occupied a relatively sleepy corner of the multitrillion-dollar artificial intelligence industry. Now they're being asked to come to its rescue.
While Anthropic and OpenAI are at the heart of a fierce debate over whether they can safeguard their advanced models and grow their businesses simultaneously, they're seeking support from a handful of small third-party groups like Model Evaluation and Threat Research (METR), Apollo Research, and Transluce.
The evaluators, which mostly operate as nonprofits, are still finding their footing in an industry where capital is flowing at historic levels and new models are rolling out faster than ever. Their primary role has been to assess AI model capabilities and risks and to flag instances where the technology behaves badly.
Without a federal push for regulation, evaluators have taken on outsized importance. Anthropic CEO Dario Amodei pledged to embed independent evaluators in his company last month – a move that OpenAI CEO Sam Altman quickly endorsed. President Donald Trump supported the idea, as did most of the largest U.S. tech companies. But left unanswered are questions about how those third parties should be funded, what level of access they will have and what the reporting structure will ultimately look like.
"To a degree, the problem, as always, is money," Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC in an interview. "Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this."