Article
AI News Agentic AI

The problem with Demis Hassibis' AlphaFold is reasoning rather than data, says new research

by TechDefused Newsroom
The image depicts a vibrant representation of a DNA double helix, showcasing colorful atoms and molecules swirling around its structure. The artistic portrayal highlights the complexity and beauty of genetic material in a visually striking manner. — Credit: Photo by MJH SHIKDER on Unsplash c Photo by MJH SHIKDER on Unsplash

The argument that AI-driven science must now run on agentic reasoning rather than on scaled data-trained models rests on a premise worth examining.

That premise is that AlphaFold was a fluke of circumstance, made possible by a corpus of standardised protein structures that most fields simply do not have and cannot expect to acquire within a useful timeframe.

The Protein Data Bank is valued in the piece at roughly $21 billion of accumulated experimental effort, assembled over decades, and treated as effectively irreproducible.

Waiting for equivalents elsewhere, the argument runs, would take decades, so the field should instead deploy generalist agents that can plan, hypothesise, critique and iterate on messy data.

The conclusion may well be right. The reasoning behind it skips the more interesting question.

What the Protein Data Bank actually was

The PDB was not a natural deposit that biology happened to sit on top of.

It began in 1971 at Brookhaven National Laboratory with seven structures, distributed on punch cards and magnetic tape by post.

The National Science Foundation started funding it in 1975. The Department of Energy and the National Institutes of Health joined in 1989. By 2003 it had become a worldwide consortium spanning American, European and Japanese institutions, and it now holds well over 230,000 structures.

What made it usable for machine learning was not the money alone but the governance around it: mandatory deposition enforced by journals and funders, a single standardised format, and unrestricted public release.

Every one of those is a policy decision. None is a law of physics.

The reason no comparable corpus is being assembled today for cell biology, materials science or immunology is not that the measurements are impossible.

It is that nobody is being paid to standardise them, deposit them and give them away.

And the demonstration effect of AlphaFold has made that outcome considerably less likely, because it established precisely who captures the value when a public commons meets a well-capitalised model builder.

DeepMind trained on data the taxpayer bought and spun out Isomorphic Labs to monetise the result.

Anybody assembling a comparable dataset now has a clear commercial reason not to publish it.

Hypotheses are the cheap part

The flagship example for the agentic alternative deserves closer reading than it usually gets.

Google's AI Co-Scientist was given a question that José Penadés and colleagues at Imperial College London had spent roughly a decade answering in the laboratory, concerning how capsid-forming phage-inducible chromosomal islands spread between bacterial species.

Within two days the system's top-ranked hypothesis matched the team's unpublished conclusion, that the islands hijack diverse phage tails to widen their host range.

That is a genuinely striking result and it is also, on Google's own account, an assistive one.

The company's write-up notes that the mechanism had already been experimentally validated in the laboratory before the system was used, and that the model was drawing on decades of prior open-access literature.

The paper describing the exercise is titled to indicate that the AI mirrored experimental science rather than superseding it.

Strip out the marketing and the case demonstrates something narrower but more useful: generating plausible hypotheses is now nearly free, while establishing which of them is true costs the same decade it always did.

Recent work on agentic self-driving laboratories makes the same point in blunter terms, observing that agents can automate ideation, planning and analysis, but that final validation still depends on real experiments.

Move the reasoning layer to near-zero cost and the constraint does not disappear. It relocates to the bench, the assay and the measurement queue.

Funding the head, not the hands

Which brings the argument round to money, where the policy direction is currently pointing the wrong way.

Amazon, Alphabet, Meta, Microsoft and Oracle are on course to spend somewhere around $800 billion on data centres and systems this year.

Over the same period the White House sought cuts of about 40% to the NIH and more than half of the NSF's budget for the 2026 financial year, before Congress refused and delivered the NIH an increase of just under 1%, to $48.7 billion.

The 2027 request comes back for roughly 13% at the NIH and 55% at the NSF.

In July the Office of Science and Technology Policy published a framework proposing a shift away from routing federal research money through universities and towards AI-driven work at the national laboratories.

That is the same bet expressed as policy: fund the reasoning, not the measuring.

If the diagnosis is that science lacks the datasets to train on, the remedy that follows is not obviously more agents.

It is the unglamorous business of paying people to make measurements and give them away, which is what worked the last time.

by TechDefused Newsroom