Complete AI Training
Sign inGet my AI kit

Your job's AI kit

Get your AI kit

Tell us who you are and what you do. We show you your kit right away and email you the link: skills, prompts, AI agents, MCP servers and courses for your job.

500+ jobs ready, and we make a kit for any other job. No payment needed to look.

Share

AI news ·

Why AlphaFold didn't solve protein folding, and what comes next

Google DeepMind and Biohub leaders say AlphaFold did not solve biology-proteins are disordered and context-dependent, and the real work of understanding disease is just beginning.

Google DeepMind's Pushmeet Kohli and Biohub's Sal Candido pushed back on the narrative that AI has solved biology's hardest problems. Speaking at a recent panel moderated by Brandon Anderson, both leaders argued that while tools like AlphaFold represent a historic advance, the real work of understanding proteins, cells, and disease is just beginning. Their discussion framed the next decade of AI-driven biology not as a victory lap, but as a disciplined search for the right data, the right problems, and a pragmatic path toward clinical impact.

The bitter lesson isn't about data volume - it's about finding the scaling law

The conversation opened with a challenge to a core AI assumption. The "bitter lesson," coined by Rich Sutton, holds that methods that scale with compute eventually win. Anderson asked whether the same logic applies to data. Candido pushed back on the simplistic version. "One misconception of scaling laws is that scaling laws are everywhere and they always exist," he said. "A lot of the work is actually finding that scaling law."

He pointed to protein language models trained on metagenomic sequences - messy, often incomplete data - as an example. That noisy data improved performance on designing real proteins. But Candido warned against the lazy path of scaling only what is easy to generate. The Biohub approach, he said, starts by asking what data the problem demands, then building the community and resources to generate it. Kohli reinforced this, describing the bitter lesson as a warning against religious devotion to either modeling or data generation. "The problem comes first, and you should be flexible in your solution space," he said.

Why AlphaFold didn't finish the job

Kohli was direct about the gap between public perception and scientific reality. "When people say the protein-folding problem has been solved, at a conceptual level there have been advances," he said. "But think about the narrative of proteins being the building blocks. Proteins aren't blocks, and they don't act as blocks."

He described proteins as disordered, context-dependent, and still poorly understood at the level of true ground-state distributions. AlphaFold succeeded at a narrower task: replicating structures deposited in the Protein Data Bank. That turned out to be useful. It did not capture dynamics, function, or the full range of conformational states. Kohli said he once tried to bypass PDB structures entirely and train directly on cryo-EM micrographs, hoping to extract richer distributional information. "I tried it, but it requires more work," he said. He believes the approach holds promise for the next generation of models.

Candido extended the critique with an analogy. Current models, he said, are like studying a single spoke to understand a bicycle. "Those models can get better and better over time, but what you really need to do is move from models of spokes to wheels to whole bicycles, because that's what people want to understand." The goal is to model proteins in their biological context - interactions, pathways, cellular environments - not in isolation.

Scientific intuition still has a seat at the table

Anderson asked whether the field should abandon handcrafted, artisanal models in favor of general-purpose scaling. Kohli said the craft in AlphaFold 2 was deliberate. "Scientific intuition from biophysics and biochemistry told us that amino acid residues are not just doing their own thing. They're influenced by other residues. So let's bake that in." That inductive bias made the model far more data-efficient.

Candido agreed but added a nuance. Inductive biases help when data is scarce. As data scales, incorrect biases can become constraints. He also pushed back on the idea that scaling lacks craft. "There's a lot of algorithmic work that goes into taking models, training them on more data, putting more compute into them, and making them bigger," he said. Architectures are evolving beyond standard transformers, and the challenge is finding the right architecture for each scale of data.

Interpretability is in the eye of the beholder

The panel took an unexpected turn on the question of understanding versus black-box utility. Kohli argued that some level of understanding is non-negotiable - specifically, calibration. If AlphaFold produced accurate structures but gave wildly miscalibrated confidence scores, "who would trust it?" he asked. Behavioral characterization - knowing what a model can and cannot do - is essential for safe use.

But he drew a sharp line between that and internal interpretability. "Interpretability asks another question: Interpretable by whom? If you're saying interpretable by a human rational system, with the cognitive and computational limitations of the human brain, then no, AlphaFold 2 is not interpretable." He floated the possibility that a future frontier model, given access to activation layers, might produce a theory of how AlphaFold works - even if humans never grasp it.

Candido pointed out that protein language models already contain more knowledge than researchers have extracted. "People know that protein language models learn some notion of structure within their representations, but we find information about functions and motions as well," he said. He sees interpretability not as an abstract exercise but as a route to unlocking scientific knowledge that has been compressed into the model through evolution.

Why this matters for science and research professionals

For researchers and product developers working at the intersection of AI and biology, the panel offered a clear set of signals. First, the era of treating biology as a solved data problem is premature. The most impactful work will go to teams that identify the specific scaling laws for their domain - not those who simply throw compute at available datasets. Second, the craft of model building still matters. Inductive biases grounded in biochemistry, when used deliberately, produce models that learn from less data and generalize better. Third, the next frontier is contextual. Moving from isolated protein structures to dynamic, multi-scale models of living systems will require new data modalities, likely including direct cryo-EM inputs and cellular context. For professionals building or funding these tools, the message is pragmatic: define the problem with precision, invest in the right data rather than the easy data, and treat model interpretability as a functional requirement - calibration and behavioral boundaries - not an academic luxury.

Share