Former Google DeepMind researchers are developing a human‑AI hybrid to oversee model safety and keep large models from producing harmful or unpredictable behaviour.
Bloomberg reported the project positions human reviewers as a proposed safety panacea but says those reviewers will need organisational support, specialised tooling and new evaluation regimes if they are to scale beyond small, expert teams to underline that institutions must design for humans and agentic systems together.
The effort arrives as labs and regulators are rethinking how safety tests are run after recent model‑related breaches and debates about moving evaluations into more realistic, and riskier, environments have intensified.
Backers plan to measure where human judgement adds the most value by building benchmarks and comparing hybrid review against AI‑only and human‑only reviewers as a next step.