Article
AI News AI Platforms Multi-agent workflows

AI safety watchdogs guarding the guards have a problem of their own

The outside firms testing frontier AI share funders, staff and worldview with the labs they scrutinise.

by TechDefused Newsroom
The image depicts a hooded figure sitting at a computer with glowing red and orange code on screens in the background. The focus is on the back of the figure, emphasizing a mysterious and potentially illicit activity related to hacking. aiImage created using AI — Midjourney

The AI industry has hit upon a tidy answer to an awkward question. Who checks whether the companies building superhuman software are being straight about the dangers? Not the government, if the labs can help it. Instead a clutch of specialist outfits now poke at new models before release, hunting for the capacity to write malware or help cook up a bioweapon.

The names are becoming familiar to anyone who reads the technical papers. METR, short for Model Evaluation and Threat Research, measures how fast models are getting at doing their own research. Apollo Research asks whether a model will scheme against its makers or clock that it is being tested. SecureBio probes the biology. Gray Swan tries to break the safeguards. The labs credit them politely in the footnotes.

It is a neat arrangement. It is also almost incestuously close.

A very small world

Consider the money. Jaan Tallinn and Dustin Moskovitz, two of the movement's deep pockets, have written cheques to Anthropic, the safety-first lab behind Claude, and to the very evaluators meant to hold such labs to account. Consider the people. The pool of machine-learning talent capable of stress-testing a frontier model is tiny, and it flows freely between poacher and gamekeeper. Critics call it a revolving door. Defenders call it the only way to get anyone qualified in the room.

Then there is the worldview. A striking share of these testers came up through effective altruism, the earnest, spreadsheet-driven philosophy of doing the most good per dollar. Shared conviction is handy for recruitment. It is less handy when everyone in the room worries about the same risks and misses the same blind spots.

The evaluators are not naive about this. Many refuse payment from the firms they assess. Some recuse staff from awkward conversations. And occasionally they bite: METR publicly disputed one of Anthropic's own risk assessments, arguing the evidence did not support the comforting conclusion.

That is the case for the defence, and it is not nothing. But the deeper objection is harder to shake. If you distrust government regulators as too slow and too clueless, and you also distrust private auditors as too cosy, you are left short of anyone credible to trust at all.

The honest answer, floated by the more thoughtful voices in the field, is that there is no perfectly independent solution, only a wider and more argumentative one. More firms. More ideologies. Perhaps even the beancounters: the likes of KPMG, the accountancy giant, dragged in to apply the dull discipline of a proper audit. It would be slower and less glamorous. It might also be the point.

by TechDefused Newsroom