Google DeepMind: Can LLMs make genuine scientific discoveries?

by

Large language models are becoming increasingly capable reasoners. They can identify patterns across huge datasets and derive complex conclusions from established premises. But genuine scientific breakthroughs often require something different: reframing the problem and proposing premises that did not previously exist.

Google DeepMind researcher Tom Zahavy separates scientific discovery into three modes of inference: induction finds patterns in data; deduction derives conclusions from rules; and abduction proposes a new explanatory framework for a surprising result.

Today's LLMs are strong at induction and rapidly improving at deduction. The paper argues that they still lack the abductive “jump” needed to originate new foundational hypotheses. Given Einstein's equivalence principle, a model might derive much of the later mathematics. The harder step is inventing that principle when observations are sparse.

Why this matters

This is not an argument that AI cannot contribute to research. It is a warning that automated proof, experimentation, and optimization are not the same as genuine scientific invention. The proposed path forward is not only larger language models, but physically consistent, multimodal world models that can act, observe consequences, and turn simulation into new formal hypotheses.

LLMs may be excellent at climbing an existing hill. Whether they can jump to a hill that did not exist before remains an open question.

Paper:

2 views

Add a comment

Replies

Be the first to comment