The Einstein test: what happens when AI tries to rediscover relativity?

Wait 5 sec.

In 1915, Albert Einstein unveiled his general theory of relativity and transformed our view of the fabric of the physical world. The theory, which explains gravitation as a deformation of space-time by mass, is a pinnacle of modern physics that underpins cosmology, from work on black holes to measurements of gravitational waves, and is used routinely to guide space missions and GPS satellites.It has also become a yardstick for leaders in the field of artificial-intelligence technology, who are asking whether their creations could ever make a breakthrough on that level. At the India AI Summit in New Delhi this February, Demis Hassabis, the co-founder of Google DeepMind in London, proposed training a large language model (LLM) on all that was known before a particular cut-off date — he suggested the year 1911 — to see whether it could reproduce general relativity. “That would be a good test for AGI,” Hassabis said, referring to the nebulous concept of artificial general intelligence that is a goal for many in the AI industry.How AI is reshaping discovery in maths and physicsSuch a test needn’t specifically involve general relativity. In December 2024, Owain Evans, a researcher at the non-profit organization Truthful AI in Berkeley, California, gave a talk about ‘vintage’ or ‘historical’ LLMs, which would be trained only on historical data up to a certain date. Evans asked what such models might be able to rediscover.This year, several teams have built instances of vintage models, including an effort at Hassabis’s test. But their early attempts reveal more about the limitations of current AI than about its strengths.A relativity-like breakthrough is not inherently out of reach for AI, says Ido Kaminer, a specialist in quantum optics at Technion —Israel Institute of Technology in Haifa, who co-authored a preprint titled ‘Can AI follow in Einstein’s footsteps?’, posted in July1.But, he and his colleagues argue, it won’t happen without rethinking some of the principles on which today’s models are built.Creative leapsHassabis had mentioned the Einstein test analogy in media interviews last year, but another Google DeepMind researcher, Tom Zahavy, described it in detail in a position paper posted on his website in January. Titled ‘LLMs can’t jump’ (see go.nature.com/3ykartg), the paper highlights the current inability of LLMs to make jumps of reasoning like Einstein’s.Zahavy wrote, as philosophers of science have long recognized, that such advances require not inductive reasoning that derives a general rule from the accumulation of data or examples, but abductive reasoning: “a creative leap that invents a cause for a singular phenomenon”.That isn’t the forte of current AI models, which seem better suited to doing the grunt work of science than to making transformative discoveries. The models look for correlations in vast data sets, through being trained on known examples and then guided by prompts to supply the most statistically likely answers for unknown cases.‘It is incredible’: How AI is transforming mathematicsThis makes them valuable predictive tools and can be helpful when there are plenty of data around. But imaginative leaps, or what twentieth-century historian of science Thomas Kuhn called paradigm shifts in understanding, often arrive when there are few data points to go on, or arise from anomalies that don’t quite fit the standard picture.The way scientists often proceed in such a situation is to develop a ‘world model’ — an inference about underlying principles — from small amounts of data. For example, careful observations of planetary motions by the seventeenth-century astronomer Johannes Kepler, and the mathematical relations that he deduced, led Isaac Newton to develop his gravitational law and his mechanical laws of motion in response to forces.Can AI models similarly reason their way from sparse data to general world models? Not currently, argues Sendhil Mullainathan, a computer scientist at the Massachusetts Institute of Technology (MIT) in Cambridge. In a conference paper published in July2, he and his colleagues provided an ‘orbital mechanics’ foundation model with synthetic data on various planetary systems that obey Newtonian mechanics.They found that the model never inferred the true law of gravitation (relating the force between two bodies to their masses and their separation), but inferred a different law for each planetary system, each wrong in a unique way.Still, LLMs have been making striking advances in mathematics, in ways that suggest they can sometimes intuit logical structures underlying their training data. A finding this May from an AI chatbot, devised by the company OpenAI, that disproved an 80-year-old conjecture by Hungarian mathematician Paul Erdős was a “genuine conceptual advance”, says Mullainathan.The AI didn’t just search its way to an answer by brute force, he says; the solution needed new, small abstractions along the way. At the same time, the disproof did bring together ideas already extant in the mathematical literature — so, in that way, it wasn’t an Einstein-like flash of discovery from seemingly nowhere.A moment in timeAfter Hassabis brought up the 1911 test, independent AI researcher Michael Hla, in San Francisco, California, tried something like it. Writing on his website in March, Hla said that he trained an LLM he calls Machina Mirabilis on pre-1900 data, to see if it could produce quantum mechanics, Einstein’s special theory of relativity (1905) and general relativity.AI isn’t ready to research itselfHla gave it helpful nudges: in one instance, he prompted it with observations about the photoelectric effect (in which light knocks electrons from metal plates), which would later be explained by Einstein’s quantization of light. In another case, he gave it the essence of the ‘elevator’ thought experiment that would lead Einstein to general relativity.In some cases, Hla argues, the model showed “glimpses of intuition”. For example, when presented with the photoelectric effect, it stated that the light “breaks up into a multitude of distinct impulses” (hinting at the idea of light quanta). But the LLM failed in most cases, lacked any true understanding of the physics it adduced, and at times was “parroting words that seem plausible”, but seemingly without “any sort of strong internal representation of the world to reason from”, Hla wrote.Hla also found that it’s very hard to train such a historically constrained LLM, when the data are riddled with broken English and artefacts from digitizing printed material.That’s also what independent computer scientist Nick Levine and his co-workers found in work presented online in April (see go.nature.com/4ilz5sp), in which they attempted to build a vintage AI model using only what was known up to 1930 — a year chosen because works published in that year entered the public domain in the United States at the start of 2026.The opening of Einstein’s handwritten manuscript on his general theory of relativity.Credit: David Silverman/GettyWith such a model, says Levine, who is based in San Francisco, in principle “we could try to test questions around things that were developed in the 1930s”, such as Turing machines (a central concept in the theory of computation), Gödel’s incompleteness theorem in the foundations of mathematics, or the particles called neutrinos (postulated in print in 1934).Levine and his colleagues discovered that it is hard to create a vintage model fed on only what was known up to 1930; the training materials are maddeningly leaky. “If you ask it about what happened in the 1950s, often it’ll just accidentally answer,” says Levine — and often correctly. The supposedly pre-1930s model, for example, answered questions on the administration of Franklin D. Roosevelt, who was US president from 1933 to 1945. Filtering out information from after a certain date is difficult when data sets aren’t precisely or accurately dated.All the same, Levine (who worked for a time as a quantitative forecaster in economics), thinks that ultimately it should be possible to ask this “1930 mind” to make forecasts, which could include “the kinds of things that you’d see in a prediction market”.Prediction-type AI models, he says, could work for forecasting scientific discovery, too. He anticipates testing whether a model trained on data up to, say, January of this year could come up with a discovery now known to have been made in June. “I would predict that a model [like this] will make a non-trivial discovery, just because of the scale of the data and resources involved.”Researchers at the University of Zurich in Switzerland have created an entire family of historical models, in a project called Ranke-4B (see go.nature.com/46efotz). These are trained on time-stamped text with historically resonant cut-offs of 1913, 1929, 1933, 1939 and 1946. Team member and economist Daniel Göttlich, now at the Swiss Federal Institute of Technology (ETH) in Zurich, says that there aren’t enough historical data or computing resources available to make these as strong as modern LLMs, so he hopes to test only whether a historical LLM might generate ideas with traces of future advances. “What we are testing for, figuratively speaking, is not so much genius, but sparks of genius,” he says.Too many theoriesThere’s no obvious obstacle to an AI language model producing, say, the general theory of relativity, argues computer scientist Jacob Andreas at MIT, as a probabilistic output among many other, entirely bogus, theories about a given body of data.The problem, he says, is that it’s hard to distinguish between theories that are correct, or at least worth testing, and those that are wrong. This was the issue, for example, with the various (erroneous) ‘gravitational laws’ generated by the orbital-mechanics model of Mullainathan and his colleagues. In mathematics, by contrast, it’s possible to immediately verify each step in an LLM’s workings as true or false.Why AI systems are most useful as designers of new scientific toolsThis reflects a distinction between the ways in which LLMs and humans devise theories: people rarely build a range of possible theories probabilistically and then go about finding which, if any, is correct.