OpenAI has neared a deal with Anthropic to stress-test each other's AI models, The Information reported.
The two are among the leading builders of frontier models, the most advanced AI systems on the market, and the talks would tie them into a structured safety partnership.
The proposed pact was described as legally binding, and built to let each lab run adversarial tests, deliberate attempts to make a model misbehave, against the other's systems under controlled conditions.
Whether the two have signed anything is not yet clear.
Not the first time
The interesting part is that the rivals have done a version of this already. In 2025 they ran a first-of-its-kind joint evaluation, testing each other's public models and stripping out certain safeguards so the systems could be pushed in a rawer state.
That exercise threw up real findings, among them higher rates of hallucination and sycophancy, the tendency to flatter a user rather than answer straight, in some of OpenAI's models.
A binding deal would turn that one-off into a standing commitment, at a point when regulators, customers and governments are pressing for independent testing of advanced models.
A rivalry with friction
Cooperation has not been frictionless.
Around the time of the joint tests, Anthropic briefly cut OpenAI's access to its Claude models, citing a breach of its terms of service.
That the two would now formalise a safety pact despite that history is what gives the story its edge.
A wider push
The idea has backers beyond the two labs.
At last week's All-In Summit, Elon Musk argued that competing AI labs should check each other's models for security flaws through a peer-review-style process before release.
The OpenAI-Anthropic talks would put that principle into a contract between the two firms best placed to test it.