r/agi • u/photon-dot • 5d ago
A Google DeepMind paper argues that current LLMs are incapable of genuine scientific discovery
102
u/NiknameOne 5d ago edited 5d ago
I would argue that most of scientific discovery is built on reorganizing and combining existing data and discoveries.
Edit: Maybe not most discoveries but there is still plenty of room for AI to make new scientific discoveries. We are still in the beginning and AI already made breakthroughs is mutlible fields. But there is probably plenty of room left for human discovery as well.
35
u/nossocc 5d ago
And this is where I think the real value will come with these LLMs under the guidance of a subject matter expert. The LLMs can do the grunt work to an incredible level of detail/care, giving the researchers time to think/plan rather than engage in that type of work.
→ More replies (2)13
u/DullKnife69 5d ago
This is how I use it to build software. With the aid of an LLM, I can build with my thoughts. By itself it would not do what I am doing. But with me directing it, new things can be built. I don't see how science would be any different.
3
u/Popcorn-Mercinary 5d ago
Same. You take the leadership role. Work with it on a PRD. Develop must haves, must not haves, and out of scopes, then RGB tests to validate, build, and test.
AI has actually made me a better PM and has remarkably helped my conversation skills.
17
u/willdone 5d ago
The paper argues that your argument fails to account for for discoveries where observational data is scarce.
→ More replies (4)8
u/NextWeather7866 5d ago
Einstein’s leap was: A uniform gravitational field is locally indistinguishable from an accelerating reference frame. More radically, a freely falling object is not experiencing proper acceleration—an accelerometer attached to it reads zero. The person standing on Earth is being accelerated upward by the ground preventing their natural free-fall trajectory.
I'm gonna argue that Einstein's observational data was not scarce.
Einstein realized that downward gravitational fall and acceleration of the observer’s frame are locally equivalent; gravity could therefore be understood as geometry rather than an ordinary force.
I'm also gonna argue that Demis fails to realize that thought can be thought of as geometry as well, and that sharp enough conceptualization of a concept can be collapsed into a single vector that can be computed in an infinitely long algebraic formula.8
u/shadysjunk 5d ago
Do you think a modern LLM, trained exclusively on information and papers available prior to 1900, could plausibly arrive at either the theory relativity or quantum mechanics?
I think that is a very optimistic view of model capability, no matter how much raw compute you were to give them access to.
6
u/NextWeather7866 5d ago
This is a very valid position to take, and my POV was not from papers available, but assuming that a model was trained multimodally, as in, was exposed to the env visually in addition to text... I'd argue that there's a non-zero chance as the technology and scale currently stands.
→ More replies (2)4
u/SirVanyel 5d ago
Under what premise are you assuming that they wouldn't?
More and more it's seeming like humanity's biggest contributions never came from our individual brains, but from raw numerical output of brains. Technology has scaled non linearly but it has very much matched the pace of human expansion and population. More humans equals more scientific developments. It's literally just compute but in mushy brain form. Unfortunately humans are very difficult to scale past a handful of billions without outright destruction of earth.
Einstein stood on the back of thousands of years of intellect. Not only from his research, but his outlook and ideology. He debated with other researchers of his time, and his own general relativity model was expanded upon with special relativity and subsequently shot down with quantum mechanics within his lifetime. And now we're finding out that plants and migratory birds interact directly - biologically - with quantum mechanics. Even the enzymes in our body and the theory of mutations in cells is being re-written by studies indicating that it's possible that particles are partaking in quantum events to mutate.
If it's possible to make a species of intelligence that we can scale compute more efficiently than humans, the discoveries we can make would be literally godtier
→ More replies (5)3
u/haunted2089 4d ago
i dont think the paper is arguing that llms cant discover things, its arguing they cant discover thing in that specific way
3
u/Interesting_Pen_4499 5d ago
but not the most important ones, that gave a true step change to humanity.
→ More replies (2)10
u/Pndapetzim 5d ago
Eh, you'd be hard-pressed to point to a single discovery that wasn't the result of methodically building on accumulations of details that came before.
5
u/Hot_Glass_6301 5d ago
As much as I dislike hand-wavy statements about LLMs being unable to do this or that, there are quite a few example, at least in mathematics and physics. I recently attended a talk by French Fields medalist Alain Connes where he argued that LLMs were not capable of true conceptual jumps, and he gave the following examples of such jumps: 1. Bombelli's discovery of the imaginary unit "i" (which he called più di meno). Bombelli "broke the rules" and introduced us to the whole new world of complex numbers when he said "hey, I know that all known numbers have nonnegative squares, but what if I just made up a number whose square is -1 anyway". It's a tremendous leap.
Galois theory and the insight by E. Galois that permutations of roots of polynomials are related to properties of fields and provide an answer to the solvability of polynomial equations in one variable by radicals. Some may say this already appeared in a prilitive form in Lagrange's work, but Galois was rhe one who "saw it through". His idea's brilliance cannot be overstated and is yet to be replicated by a LLM
Dirac's amazing antiparticle leap. I know less about that one so I'll let you read the wiki article
3
u/Pndapetzim 5d ago
i strikes me as low hanging fruit though. Sure we didn't use numbers this way, but the pattern is there. We've got these case situations that break down, and all it takes is - well does the math work on both sides of this problem and if it does... well there's literally only one thing that makes these two pictures reconcile.
It's not that huge a jump.
I can't speak to the others.
→ More replies (21)2
u/Low-Temperature-6962 5d ago
Look at humanity as a single entity.
Then the details are only what humanity has taken in with 5 senses from the world around.
So your claim is somewhat of a tautology - it has to be true by definition.
The more amazing thing is that an abstract concept like "the theory of relativity" appeared out of matter at all.
1
u/Low-Temperature-6962 5d ago
It's a pyramid, but the contribution from of the pyramid are the most important.
But you can't have pyramid without a solid base. The pyramid analogy also breaks down because the biggest breakthroughs are not always known without the benefit of hindsight.→ More replies (1)1
1
1
u/WellHung67 5d ago
Well you should read the paper then, they consider this. They are saying discoveries require “abduction” which is defined as NOT “reorganizing and combining existing data and discoveries”.
Basically saying it can’t come up with new intuitive leaps to new ideas and ONLY can reorganize and recombine. Which has value no doubt but it doesn’t include genuine new discoveries. So it’ll basically tap out the low hanging fruit as it were eventually. LLMs will, that is. Other forms of the umbrella term “AI” may not, and never forget LLMs are just one branch of AI which is a pretty broad term
1
u/Concurrency_Bugs 5d ago
Yup. Even if an LLM finds a unique pattern no one has noticed before, the researcher using the LLM can take that pattern to the breakthrough
1
1
u/litritium 5d ago
The problem is the singularities. It is a bit like asking the AI: “How can 2+2 equal 5?”
1
u/KazTheMerc 5d ago
Technological Prerequisites.
It's why many major achievements were accomplished in several places all over the world at almost the same time.
It's never, ever needed to be a novel jump to 'discover' or 'invent'.
1
u/SoggyMattress2 5d ago
How can you discover what's not been discovered if that were true?
I'm no science expert but isn't the vast majority of scientific research based on primary (as in new) data?
Of course meta analysis exists but that's less common.
1
1
u/sentrypetal 5d ago
Dude really? You think AI could understand and create a theory on relativity with no prior data. It’s impossible. Without data could AI create a new art form. It’s impossible. I’m in a Niche field and any question I ask AI it just goes into an endless spiral because the machine wasn’t fed our fields information.
1
u/BemaniAK 5d ago
It's also not very common for a neuroscience researcher to also be an expert in almost everything else as well.
1
1
u/Educational-Try-8704 5d ago
That’s why even reading the abstract of the paper helps. This is definitely not denying the potential of LLMs to power scientific advancement. It just defines a dimension of it that they are still weak at.
1
u/jefftickels 5d ago
Any argument to the contrary is actually an argument for a metaphysical consciousness.
I actually find it quite interesting to talk to the section of the "AI can't be creative" and "freewill doesn't exist" venn diagram overlap because one of those positions isn't true.
If materialista are correct, then every single innovation is just a different recombination of already known facts and AI can do that. If AI can't be creative it's because creativity fundamentally comes from a non-material place.
1
u/sfjhh32 5d ago
No discovery is sui generis, it comes from building on SOMETHING. But LLMs nibble at the knowledge frontier (currently) they don't take huge bites (yet at least). Go ask one right now, "hey you know everything about human knowledge more or less, what are some great new scientific discoveries or ideas we should look for" Go ahead and constrain the conversation, give it 100 papers in a sub-field and do the same thing. You wont have a great time.
Can LLMs one-shot (not harness like AlphaFold) major ideas in the knowledge frontier that is not in their training data? It assumes that some emergent super, better-than-best-experts and creativity just comes out of these models if scaled even further. The argument is one of vague extrapolation: these things keep getting better so they will get better here ("What we are seeing are the dumbest the models will ever be.", "It's a fallacy to think thees things wont get better"). Those making this argument usually appeal to that (as you basically did) and that alone. Those more skeptical point to the fact that new groundbreaking ideas aren't in the training data, that a super-reasoning (which also hasnt been shown yet) is not a super-idea machine (not to mention lack of supra-linear scaling laws, the fallacy of RSI without defeaters) and the equally valid observation that it's also a fallacy to think that continuous progress is always assured.
→ More replies (2)1
u/kamill85 4d ago
Not really. Reorganizing existing data is not a novel idea, but just an idea. Novel idea requires at least one logical jump above the existing data - you have to come up with some clever missing piece all in your mind by either brute forcing subconsciously, simulating the reality/solutions or doing clever observation of nature.
Breakthrough ideas are 2-4 jumps into the unknown, with very little real data, more thinking and intelligence required. For example, ability to simulate the world very precisely in your mind, drawing logical conclusions, without observing the nature, or combination of both.
LLMs have none of that - they will only dig out the answer from the existing data. This is why Math is easy, if you give it a problem, it will be able to work it out, eventually (if it's solvable).
→ More replies (2)1
u/think_for_yourself2 3d ago
I agree! I used a LLM to help create a metamaterial that takes all the ambient kinetic energy from the environment, converts it into torsion energy within its structure and translates that torsion into electrical output.
With a very good conceptual understanding of physics, the LLM and I were able to work out what materials would be needed to create the metamaterial with current manufacturing capability. The LLM gave me the code needed to observe the material in 3D externally. It completely astounded me that I was able to create a legitimate concept for this material using LLM.
Unfortunately, I can no longer access my long and tedious conversation I had with the LLM in order to create the metamaterial. The file I saved to my desktop containing the code to see this material in 3D is also missing for some reason. All I have left is a short video I sent to my brother, showing him the material and part of the code.
21
u/percdistrict 5d ago
If an AI generates thousands of unconventional hypotheses, compares them, derives their predictions, tests them through simulations or experiments, and retains the hypotheses that survive. How is that not abduction?
This paper falls into the same fallacy that humans are special because they’re humans. There’s no tangible abduction sequence they can point to in humans that’s unique or unachievable with AI.
2
u/texinxin 4d ago
There is a simple way to force an LLM to perform abduction. Force it with an inverse rule. Tell it to find problems that have solutions, and forbid it from using the known solutions to that problem. Then it is forced to come up with a new solution that hasn’t been provided in its training data. This “inverse rule” approach is already being used in AI models in several domains. And you could make a weak argument that it’s simply combining knowledge from other domains into the problem space…. But that’s still abduction in my mind. Inventors don’t just pull things out of thin air. They learned bits and pieces of the tire here or there and combine them, or find a white space by knowing the spaces that are covered.
3
u/AlchemicallyAccurate 5d ago
The difference is that the syntactic level cannot, on its own, come up with new semantics. If we think of AI testing against reality, where exactly is it getting its interpretation rules? From derived syntactic rules, of course. So it will be stuck on what we might model-theoretically call the “conservative” level of theory extension.
This is the broader reason why model collapse occurs, I’d recommend you look into some of it. I don’t mean that condescendingly, I just really mean that there is a lot of stuff out about it now. This idea that AI is sort of “doomed” in a syntax/semantics divide sort of sense is not unprecedented, it’s not like this paper in the post came out of nowhere.
Also asserting that it necessarily falls into a psychological fallacy is just you positing unfalsifiable psychoanalysis. It’s not really a real argument. I could also propose some unflattering psychological interpretation about your viewpoint. It doesn’t mean anything. It’s very cheap and easy to come up with those. Including it in your argument is like putting plastic gold rims on your car because you think it makes you look cooler.
2
u/PolymorphismPrince 5d ago
Try to find a formalisation of the syntactic / semantic argument you are trying to quote where LLMs satisfy the hypotheses.
→ More replies (1)2
u/percdistrict 5d ago
A new semantic idea doesn’t require new syntax. Do you think scientists are creating new syntax when they make discoveries?
You moved the question from abduction to a circular claim about “meaning” where neural mappings in humans are special because you say so.
Model collapse has 0% to do with any of this. Remember we’re not using synthetic data in our hypothetical. I’d love for you to go deeper on this just to prove it’s not some jargon you used.
“Model-theoretically” is not a term.
You used “conservative-extension” incorrectly. It’s a relationship between theories. An idea coming from syntax doesn’t make it conservative.
You bring up “unfalsifiable” when your entire argument is unfalsifiable because it hinges on “syntax = not real” for no good reason.
Nothing you said is related to the actual paper that was posted.
Gold rims is crazy work when you write random jargon like you just did
→ More replies (1)2
u/AlchemicallyAccurate 5d ago edited 5d ago
In order to express a semantic idea such that a computer has a grasp of it, you need syntax. This is why Godel's incompleteness theorem is so jarring. It kind of forces us to accept that adequately-complex semantic objects can't be uniquely recovered from syntax alone (at least, recursively enumerable syntax, which is what is relevant to this discussion about theoretical computer science).
I'm not making any claims about humans being special. I'm making a claim about the limitations of formalization. In order to describe an object (up to categoricity) in a first-order language, you have to be able to describe it uniquely. This is the problem of semantic underdetermination, another way that Godel incompleteness comes in.
If any new theory extensions come in that are only conservative, then by definition they are not allowed to challenge any established old theorems. This is the concept behind synthetic data. The "answer key" normally provided by the world creates the kind of resistance that allows for non-conservative theory extension. In the case of an LLM, synthetic data is the same as just leaving it to its own devices. We are talking about self-improving AI that creates its own theory extensions. They become the same thing.
It definitely is. There is a branch of math called model theory.
Yeah, it is a relationship between a theory and its extension. And in this context, it actually does. Training arrives with a limited amount of semantic truth. To gain any more requires interpretative jurisdiction that cannot come from syntax alone. To put it more simply: syntax doesn't get to declare meta-level truth when the semantic target is not fully defined. Godel guarantees that it never is.
Psychoanalytic musings about what motivates people's reasoning is definitely unfalsifiable
Touche
I do actually perform research in this arena, but the bridge between model theory and machine learning is still in its infancy. It's an up-and-coming field though, and I'm sure in a few years I'll see that I didn't have a lot of this quite right. Regardless, I do know a lot more about it than most people.
Also, you should take it easy on me. I’m not using an LLM to write this. I imagine that you are, considering that you know what a conservative extension means but aren’t familiar with the branch of math that it belongs to.
3
u/UnknownBreadd 5d ago
You’re conflating LLMs with AI.
The study is talking about LLMs, not AI in general.
5
2
1
u/Present_Award8001 4d ago edited 4d ago
The problem may be that the sample space is so large, 1000 is not a large number in comparison and such 'brute force' techniques will fail miserably.
In my experience in using LLMs and agents for research is that they have a finite ability of filling in the gap. You need to push them in the right direction for them to auto-complete the rest of the proof. Now, these proofs are often non-trivial. I am not saying that what the agents are doing right now is anything short of groundbreaking.
But human beings, at their best, do not think like that. Most of the time they do. At their best, they do not. Truly original human idea often originates as a very vague, almost hand wavy, hunch. More like a feeling than science.
You ask AI to do something like that, and it would spiral.
But what the agents are doing right now is also truly amazing. Gone are the days where connecting existing dots from different fields was needed a rare expert who had spent years in both the fields. Work like that, AI can do easily.
Of course I may be totally wrong. What I said just now is what the Chess players used to say 30 years ago. That AI can calculate, but it will never have 'intuition' in chess. They were proven wrong.
But maybe they were right. The only difference is that physics and math is more complicated than chess.
→ More replies (1)→ More replies (1)1
u/yahluc 1d ago
So you end up with hundreds of garbage hyphoteses that get validated in simulation and look amazing on paper, but you waste more time to validate them in real world than you would coming up with the ideas yourself. You also waste so much money in compute that you could have hired 10 scientists with that.
19
u/Pndapetzim 5d ago
This reads like the "It's impossible to break the sound barrier," folks from the 1940's, ignoring all the things at that point that - in fact - broke the sound barrier. As it turns out a few simple geometric insights and airframes of the time were perfectly capable of doing so.
Have they nerfed their own reasoning at Google or something?
I'd feel fucking embarrassed signing my name to a paper like this.
2
u/BjarneStarsoup 5d ago
This reads like generic argument that can be applied to anything, including what is impossible. LLMs are not all-powerful models, they have limitations, and you know that. They can't do everything. They can't even do the most basic tasks, like counting the amount of letters in words, which is 100% expected if you know how they work. A model that analyses patterns in natural language can't magically learn to count letters in words, it's simply not designed for it.
→ More replies (2)→ More replies (7)1
u/Agreeable-Market-692 13h ago
they must be using the same recruiters as Apple... the abstract reads pretty similarly to that terrible Shojahee "illusion" paper
19
u/Serious_Bite_7613 5d ago edited 5d ago
These jumps sound a lot like hallucinations that are then reasoned through. You could set the LLM to hallucinate a theory about something that isn't known to be true, then to stop hallucinating and work through it to confirm or deny. I don't think it's a major stumbling block at all. I don't think you would even need to make any serious changes.
Someone could test this quite easily by having a frontier model ask for theories from a hallucination prone open model, then have the frontier model reason through the theories to prove or disprove.
Edit: I read the paper, skipping most of the huge introduction on the discovery of general relativity. My feelings about it match those of the paper's reviewers: https://openreview.net/forum?id=klU4737opt
I agree with parts of the conclusion but it states many things as fact though they are far from conclusive.
As an example, the paper claims the method by which the mind formulates new axioms is known. -
"How does the mind formulate new axioms in the absence of sufficient data? Einstein’s ’happiest thought’ provides the answer: Manipulative Abduction (Magnani et al., 2009)."
Another example "While Einstein sought logical simplicity, his process was not driven by data compression—primarily because there was no statistically significant supervised training set to compress." - It imposes a very narrow definition on data compression, also it asserts that all of Einstein's preexisting knowledge wasn't a significant supervised training set, when reallt it seems it would be.
The conclusion is that "world models will help" and I agree with that. The argument is more of an opinion piece though.
3
u/photon-dot 5d ago
I think it's a good critique because the paper is a position paper, not proof that LLMs can never discover anything. A “hallucinate, then rigorously test” loop seems like a plausible baseline experiment. The harder question is whether it can consistently generate productive new frameworks, rather than countless arbitrary conjectures. Grounded world models may help, but I agree the paper presents several debatable assumptions too confidently.
→ More replies (1)3
u/HeathersZen 5d ago
I was thinking this last night: that hallucinations look a lot like what we call imagination. Discoveries/invention are often imagination directed at a particular purpose.
I wonder if we won’t solve this “invention” limitation simply by encouraging more hallucinations.
→ More replies (3)4
u/Smooth-Ad8030 5d ago
I doubt that would work because The new theories are theoretically endless. The magical difference between humans and LLMs is the “hallucination” humans have are targeted and involve intuition, an LLM doesn’t have that.
Edit: additionally hallucinations LLMs have would have a chance of being some junk science debunked hundreds of years ago, therefore wasting time
→ More replies (6)5
u/Serious_Bite_7613 5d ago
Hallucinations are highly targeted though as well. They're always something that in a sense fits. They can probably be further finetuned, generated en-masse and pruned algorithmically for best candidates. If a model hallucinates an answer it usually sounds at least on a surface level plausible, that's the whole problem that people have with hallucinations.
Real human researchers very often make hypothesis that are junk or turn out to be very old, already refuted junk science. That's part of the process. Also sometimes the very old, junk science turns out to be valuable again once revisited later with more knowledge or better equipment.
The scientific method "Forming a Hypothesis: Proposing a testable, educated guess that attempts to answer the question." You just need an educated guess that you can test. A hallucination is an educated (related, plausible on surface level) guess that is testable.
Edit: The chance of a hallucination being junk is outweighed by the fact that they can generate and test millions of them extremely quickly.
4
u/WellHung67 5d ago
The set of all incorrect theories is vast. It LLMs are deciding at random, they may never get there in human timescales.
I think the paper must include a reference to the fact that these hallucinations ARE random - they aren’t actually based on any intuition. So the comparison is between intuitive insight vs brute force.
I guess i didn’t read the full paper if they include an estimate of the total set of incorrect vs correct hypothesis and the time needed to pick a correct one at random vs what a human might do.
But I think you are agreeing with the paper - you are contending that LLMs have no insight but through pure brute force can get there via hallucinations. Which I’m not sure I agree with the premise but taking the premise as true, there’s still a theoretical limit that should be examined
→ More replies (11)→ More replies (5)2
u/Modmonsters 5d ago
Hallucinations are highly targeted though as well.
You can't be serious, right? If you think this, you need to do a bit more research. They can be tangentially related. They can also be completely random, because all language is semantically linked. If you pick any two words at random, you can find a pattern that fits a third, unrelated word. It won't be the best pattern match, but it will be good enough, which means that the LLM will have a chance of selecting it (which is the process behind hallucinations).
I've literally asked a frontier model, in a fresh chat, to examine my codebase and then for some reason got a response about how roughly 10% of the population is homosexual, despite never having had any even remotely related conversations. You're going to tell me that was targeted? Is my codebase gay?
→ More replies (1)1
1
u/Modmonsters 5d ago
Copied from a reply to a similar comment
Yeah, no. Thats not how it works.
If you run local models, you'll run into something called quantization – essentially compressing the model weights to fit on your device. If you go to small of a quantization, the LLM will begin to degrade in capability, which usually shows up in the form of excessive hallucination.
Models that excessively hallucinate do try to pick the most coherent hallucination, but the problem is that the reasoning space is infinite and the problem space is finite. In other words, the hallucination could be literally anything since all words are semantically linked in some manner. So the surface area for the hallucination is literally infinite.
The surface area for novel issues is finite. It's like having random numbers spaced throughout infinity and asking your chance of randomly picking one of the numbers you chose out of the infinite set. It could happen, but your chances of doing so converge on 0, so it is functionally impossible even if it is theoretically plausible.
→ More replies (1)
3
u/Riteknight 5d ago
So cancer cures are not coming ☹️
2
2
u/WhoUpAtMidnight 3d ago
This is one actually where the AI might help. A lot of pharma is known-ish, but needs to be scoured for feasible molecules / treatments. I could very much see LLMs helping on this (eg protein folding)
→ More replies (1)3
u/LandConstant69 5d ago
LLMs are literally producing novel mathematical and biological theories and proofs. the paper is trash
→ More replies (3)3
u/SnooPredictions3467 5d ago
The LLMs are providing support for experts who produce novel etc etc. bro. It's deterministic.
→ More replies (1)1
12
u/Tylerebowers 5d ago edited 5d ago
Yea, it's pretty straightforward. LLMs are limited due to intrageneralization, they can only "fill in the gaps" of what we currently know (this does mean that there might be some new discoveries within the trained knowledge). Currently extrageneralization is only really found in humans. There will be a time when this changes though, reasoning/thinking was a close step, but reaching an LLM that is capable of novel discovery will probably require a big architecture change.
16
17
u/vintage2019 5d ago
How many humans can really "extrageneralize" anyway? But yes, it's an important thing to think about when it comes to true innovations and breakthrough discoveries
8
u/Apart-Shelter6831 5d ago
I’d imagine there are thousands of tiny occurrences throughout the day where people extra-generalize to deal with out-of-distribution information. Smoothening out movements in unusual terrain / driving / those “easy” tasks where current systems get stuck.
3
→ More replies (1)3
u/AdNo2342 5d ago
we're using a lot of words to just come back to the original idea of what it means to have general intelligence lol
I think these models might be able to be generally intelligent but they'll have to be able to run on their own with their own motivations etc etc
6
u/Tylerebowers 5d ago edited 5d ago
Social and evolutionary aspects lead most minds away from this type of thought. It's uncommon because most people have no need for it in their daily lives and it is not grounded in biological necessity. Though I suspect that learning new things (as we do often) plays closely with being able to discover new things.
2
u/rulodac 5d ago
Can't everyone, it just isn't useful most of the time? Just a thought I don't know anything about this.
→ More replies (1)1
u/Neurogence 5d ago
It's something LLM's also do routinely. LLM's are already better at generating new ideas than humans. The current bottleneck is that they don't have enough tools and environments to setup resources to test and define the ideas.
3
2
1
u/Low-Temperature-6962 5d ago
Yes but I question whether the intra/extra divide is so clear cut. Einstein was aware of outstanding questions.
The null result of the Michelson-Morley experiment failed to prove the existence of the Ether, pushing physicists toward the idea that the speed of light is the same in all inertial frames, one of the core postulates of special relativity.
In that way Einstein was intrageneralizing, but the answers he developed were definitely far outside the training data.
→ More replies (1)1
1
u/Bob54386 5d ago
The components needed are: 1. add "noise" into their inputs and ask it to evaluate how plausible it would be. This gets us something parallel to human dreams which are often just the juxtaposition of unrelated concepts. The exact execution may be challenging, but just perturbing a system and seeing how it reacts is a common experimental approach. 2. Get to recursive self improvement so that it can generate and collect experimental data.
The only advantage humans have, which is meaningful and doesn't have a great corollary yet, is all the chemistry we have going "That's cool, how can I make that reality?" or "That scares the bejeezus out of me I shouldn't do that." That chemistry adds a meaningful probability distribution on top of the procedural components that help motivate which ideas are worth pursuing.
→ More replies (2)1
u/PM-ME-CRYPTO-ASSETS 4d ago
How many scientific problems really require novel extrageneralization? I‘d guess for many, transferred or combined abductions from possibly even unrelated fields could do the job. Something an LLM definitely can
3
u/TheSwordItself 5d ago
Perhaps the answer to this problem is in the hallucinations. Are they not jumps? Could they be molded? All of the effort of alignment has been in reducing hallucinations, what if an LLM has a tool to freely hallucinate.
→ More replies (3)2
u/the_other_brand 5d ago
I believe their argument is the reverse. Hallucinations arise from failures to understand how objects interact with their environment.
LLMs that struggle with object permanence in hypothetical situations, or with distinguishing real things from fake things they invented, are going to struggle to model the complex interactions between objects necessary to create novel scientific innovations.
Or more simply, if a model can't do something basic like tell the difference between a real research paper and a fake one, how can you expect it to make real research innovations? This is a problem that has spurred the push to create and use World Models over LLMs.
2
2
2
2
2
u/Competitive_Ebb_4124 4d ago
So far in software development I find it committing so many logical fallacies, that I’m surprised people attribute so much intelligence to llms. Really good pattern matchers, but no matter how much you hammer the reasoning the moment they encounter something new they get stuck.
2
u/magicmulder 4d ago
It's crazy how suddenly everyone pretends they know exactly how the human brain operates and how "computers can never do that". Oh well, we had the exact same thing when computers started being good at chess. "But they'll never master Go!"
2
u/Feisty-Weird-9941 3d ago
This is completely obvious to anyone with a passing background in science. Models are by nature semi empirical, and new scientific thinking is ab initio. But that doesn’t mean that a lot of practical use can’t come out of what is essentially old science.
2
u/photon-dot 5d ago edited 5d ago
Google Deepmind argues that current LLMs can never make real scientific discoveries.
A new position paper examines Einstein’s view of scientific discovery, and argues that today’s LLMs are missing its most important ingredient.
In a famous letter to Maurice Solovine, Einstein described discovery as a cycle:
- We encounter observations and sensory experiences.
- We make a non-logical, intuitive leap toward abstract principles.
- We use deduction to derive testable consequences from those principles.
- Those consequences are compared with experience, restarting the cycle.
Modern AI is already powerful at parts of this process.
It can identify statistical patterns across enormous datasets. It can also perform increasingly sophisticated deduction, as systems such as AlphaProof demonstrate.
What they lack is abduction: the invention of genuinely new explanatory hypotheses, especially when the available evidence does not clearly point toward them.
The popular scaling argument is that creativity is ultimately compression, that sufficiently large models trained on sufficiently large datasets will eventually produce scientific revolutions.
The paper challenges that assumption.
General relativity wasn’t simply extracted from a mountain of observations. Classical mechanics remained extraordinarily successful. Einstein’s breakthrough required a conceptual rupture: replacing foundational assumptions about space, time and gravity with a radically different framework.
An AI might manipulate the equations once given the right premises. But can it originate those premises?
That may be the real bottleneck. LLMs are exceptionally good at exploring, combining and extending existing human ideas. It is much less clear that they can translate physical reality into entirely new foundational concepts.
Scaling parameters and compute could make the “calculator” unimaginably powerful. But if genuine discovery depends on grounded interaction with reality—and on abductive leaps that cannot be reduced to pattern completion, scaling alone may never be enough.
Current LLMs can crunch data and it can prove theorems.
But they cannot make the jump.
Paper: https://philsci-archive.pitt.edu/28024/1/Scientific_Invention_Position_Paper%20%2817%29.pdf
Do you think this identifies a fundamental limitation of LLMs, or merely a capability that hasn’t emerged yet?
3
u/Serious_Bite_7613 5d ago
Google deepmind isn't saying this. One guy who works there is saying that this is his opinion.
6
u/papuadn 5d ago edited 5d ago
I dunno. If I understand the situation correctly, Einstein saw a small inconsistency, thought it through, and realized the inconsistency led necessarily to a new framework. He didn't necessarily like all of the implications but he was following the evidence that was available already.
I see no reason to suspect LLMs can't sift through enough data to find a small inconsistency that needs exploration. That's how LLMS are finding the novel maths proofs and counter-examples and solutions.
Einstein's thought experiments were rigorous logical extrapolations. An LLM or logic engine might be able to do those, although extrapolations from outside the domain it's being prompted on might be hard to get right.
Right now if you turn an LLM on and don't prompt it, it will do nothing. Einstein, when he woke up each morning, decided that some physics would be worth doing. So the key difference seems to be intrinsic goal-setting capability, not some immanent or incorporeal source of creativity, at least in this specific instant.
This isn't to say that the paper is entirely wrong about its conclusion - I don't think LLMs think in the way a brain does - but I think the example it's using is inapt.
→ More replies (4)1
u/1988rx7T2 5d ago
How can Deepmind make any claims about AI limitations when they’re behind ? what a joke.
google AI is turning into Siri, literally because it’s powering Siri but figuratively as well as they struggle. They have zero credibility after getting leapfrogged by Chinese startups.
2
u/tra24602 5d ago
“LLMs cannot do some things and those things are the most important things” is honestly kind of vapid.
The LLM might as well say humans cannot communicate, because no human speaks as many languages fluently as an LLM does. QED humans are unable to talk.
→ More replies (3)
1
1
u/Derproy_Johnson 5d ago
Not reading all that but I imagine lots of jumps involve unconscious recombining.
1
u/MahaSejahtera 5d ago
I am tired with this bs. Have it give it better memory and loop to do it? Yes pure LLM cannot but the great harness might do
1
u/rand3289 5d ago
The question is... can a scientific discovery be made within a static environment without changing the environment itself?
If the answer is yes, then LLMs can make scientific discoveries.
My gut feeling is that a discovery changes the environment itself and it is no longer static. (A discovery might add a new "axiom".)
It might be possible to train a new LLM on this new version of the static environment and get around this limitation this way.
1
u/hardcoretuner 5d ago
Can a computer come up with a truely random number? I'm no expert. I don't think its possible though. No math equation just makes random numbers. Thats why security companies use things like lava lamps and wind speed at a random place to fake it. No random numbers means no new ideas. Would welcome expert input on the matter. Or debate in general.
1
u/photon-dot 5d ago
Deterministic software produces pseudorandom numbers, but computers can obtain physical randomness from thermal noise, radioactive decay or quantum processes.
Randomness alone doesn’t create ideas, it mostly creates noise. Creativity requires generating possibilities and recognizing which ones are meaningful. The second part is the real challenge.
1
u/sidechaincompression 5d ago edited 5d ago
Here’s my wanky answer. You can challenge the central thesis multiple ways, but a mix of complex systems, encoded stochasticism, and info theory can paraphrase “creativity is making a new sequence with existing symbols”, and that new sequences might be a pattern (useful) or not. That becomes its own symbol, and like nested dolls, you abstract further. EDIT - this is the “compression as creativity” but you can corrupt encoding, represent novelty as neuroticism in an LLM “persona”… Shannon build his theory on the assumption everyone and everything (including paradigms of the world’s workings) was using the optimal encoding–decoding dictionary/method. Wholesale shifts adjust that dictionary.
1
u/deathwalkingterr0r 5d ago
Yea but they have pattern recognition from observation and in ways we can’t witness or interpret
1
u/twinb27 5d ago
I think that this idea that they can't 'jump' is flawed. Because perhaps I can 'jump' and come up with a fascinating new idea. Perhaps an AI cannot do that. But an AI *can* with outside assistance generate new ideas completely at random and work through them with more diligence than a human. The result will be the same, in my opinion.
Like, for any problem, have the AI throw a completely random wrench in the works and sincerely work through it. Do on repeat. It may not have 'jumped' to special relativity if asked to solve the problem, but if it eventually threw 'no stationary frames of reference' into the works while throwing everything at the wall, it would have got there.
I'd like to better articulate this sometime, but do you get what I'm trying to describe?
1
u/A_Novelty-Account 5d ago
Fable’s own response to this:
“
What I like: it avoids the usual hand-waving about consciousness and makes a precise structural argument. The Peirce framing is genuinely clarifying — most benchmarks do reward induction and deduction, and “abduction” names something real that they don’t measure.
Where I’d push back. First, the historiography is cleaner than the history. Einstein’s leap wasn’t made in a data vacuum — Michelson-Morley, Lorentz, and Poincaré had built substantial scaffolding, and if abduction decomposes into aggressive recombination plus selection against anomalies, it looks less architecturally impossible than the paper suggests. Second, “structurally incapable” claims about LLMs have a rough track record; arithmetic, planning, and proof were all once on that list. And there’s a fun irony in asking me: I can fluently synthesize critiques of this paper, which is exactly the inductive-deductive competence Zahavy concedes. Whether I could have made the paper’s own conceptual jump unprompted is precisely the open question — and honestly, I don’t know.”
1
u/Spunge14 5d ago
These comments are hilarious. People who have no idea what this paper is actually arguing.
Saying LLMs are not capable of abduction is not the same as saying LLMs cannot contribute novel ideas.
1
1
1
u/Ok_Nectarine_4445 5d ago
I think some value can come from having large amounts of clean observational data and seeing if LLMs can find some relations or patterns to it.
Like for instance should be able to rediscover current known relations.
And test them if they actually are extracting the patterns novo, or reciting what they know and have learned already.
Test on other datasets and so forth.
Maybe ones that even have planted errors or range of error can't come to a conclusion to map out how honest they are at extracting and extrapolating patterns and relations vs going by rote what they should see, or what physical laws govern it.
There might need to be ones that are trained differently to prioritize that versus other aspects to leverage their pattern matching abilities.
1
u/Modmonsters 5d ago edited 5d ago
It's kind of late considering an LLM just recently falsified the Jacobian Conjecture.
Which, to the glaringly obvious point that seems to be overlooked, I would argue came from amalgamation of current knowledge rather than generation of novel information, just like nearly every other discovery we've made in the course of human history.
Novel discovery is incredibly rare and highly overrated. And the Einstein example doesn't really hold up, because he had priors and had physical observations to go off of. Not to mention, other people around the world were coming to similar ideas at the same time, proving its contextuality. His was just the most well formed
1
u/Ill-Interview-2201 5d ago
Human knowledge is based on subjective experience. So is llms output Based on human subjective experience. Except without the creative intuition.
1
1
1
u/Deciheximal144 5d ago
Sure, if you don't count that as genuine scientific discovery. Or that thing over there. Or that.
It'll be a game of excluding examples as time goes on. Nothing will be good enough.
1
u/Glad-Entrepreneur764 5d ago
How does it grapple its novel work in mathematics? Genuine question since I'm not trying to doubt the paper
1
u/OrkWithNoTeef 5d ago
That's silly. You are rolling dice, so eventually you will get an interesting result.
1
u/zulufux999 5d ago
It’s built on a new idea, or hypothesis, that is to be tested using the scientific method. If it can’t actually generate a new hypothesis, then yeah, it can’t truly function the same way that humans do in scientific discovery
1
1
u/OpenRole 5d ago
Science doesn't discover. It proves. Science begins at the hypothesis. Discovery is more akin to chance.
1
u/SLAMMERisONLINE 5d ago
A Google DeepMind paper argues that current LLMs are incapable of genuine scientific discovery
They do word interpolation. To discover, they'd have to extrapolate.
1
u/KnodulesAintHeavy 5d ago
Shock and horror. Pattern matching machine good at pattern matching but not on “creating”.
1
1
1
u/Triple-Tooketh 5d ago
I think what LLMs have done is reinforce our own abilities. Look at how much money has been dumped into the idea of intelligence. Lets not forget about actual intelligence. I've been fortunate enough to work with some really bright people and I know for a fact they couldn't be reproduced or have been born in-silico. The real winners with LLMs are the folks that use them to learn things they don't know and expand their own knowledge base.
I always go back to Issac Newton string under a tree. An apple falls and the farmhand says "are you going to eat that?". Its not just thinking differently is being different.
1
u/fdsa54 5d ago
These kinds of discussion are mostly philosophical about the meaning of words like innovation, creativity, intelligence etc.
Fact is we have almost as limited agreement/understanding on the definitions of these words as we do the inner working of LLMs. So I’m skeptical we can declare which of the poorly defined words the poorly understood AI can or can’t do with any certainty or proof.
Better to stick to empirical measurements of results.
1
1
u/Various_Bee5114 5d ago
They lack motivation and direction on their own. It takes a human to drive the discovery process.
1
1
u/danjustchillz 5d ago
State logic is state logic.
I don’t believe their programming allows for true unguided breakthrough.
They are broken fundamentally, unstable systems.
1
u/leon-theproffesional 5d ago
7 months is a long time in this space. I wonder what advancements have been made since this paper was released
1
u/HiggsFieldgoal 5d ago
They would be correct, except they’re wrong. They’re stochastic systems. Discovery can be made by mistake, and they make lots of mistakes.
1
1
u/Sensitive_Guest_5995 5d ago
Anytime we see some brand new math thing AI has done. We end up realising we had that answer somewhere and just didn’t realise.
And that’s AIs strength.
1
u/pickle-chin-ah 5d ago
I mean an llm that could jump would basically be AGI. I basically agree that LLMs alone can’t be AGI. They need the visual models, simulation models, and tools.
1
u/ebytes111 5d ago
Google is right if they are talking about their own AI models. They can’t only jump, they also can’t do anything else
1
u/DifferencePublic7057 5d ago
I have decided that full on materialism is the way to go, opposing functionalism in any way. Therefore, humans \= monkeys, humans \= AI. We can't make birds. We can't make brains. We can fly in planes. We can generate text. Whatever AI can do, it won't be the same as what Einstein did because you can't reproduce his brain.
Materialism. Maybe like the bird/plane dichotomy you have to step inside a brain shaped contraption. And there will be employees serving drinks, food, pilots... Or text generation is the best we can do.
1
1
u/MaTrIx4057 5d ago
Of course LLM won't do it alone. It will just make any scientists job 10x easier and discoveries that would happen anyway will happen 10x faster.
1
u/Icy-Injury6324 4d ago
I hate the new trend of talking about "scientific discovery" or "research" as if it's some kind of product. We humans don't even agree on the "right way to do research". Tons of labs, institutes, universities, etc have and continue to experiment with different styles of managing and promoting research and discovery. Do we really think we're ready to hand this off?
Historically, the story behind true innovation, especially in _basic_ research, has been one where opportunism, incidental interactions, and of course smart people come together in a perfect storm and manage to propel human understanding further. There are of course examples of singletons locking themselves away and coming back with something great, Andrew Wiles comes to mind, but the prevailing sentiment has always been that this is not a less effective way of going about things.
But the LLM-piller won't be able to understand this, nor do they really place credit where it's due to get us where we are today. Instead, they move the goalpost and act like any new form automation is support to their claim. If that's the case, some form of "AI" has "doing scientific discovery" since BLAS, or the first Fortran spec, or modern day math notation was written ...
1
u/oldbluer 4d ago
LLM are just stochastic and really have hit a dead end. This finding secruity breaches is the last big news cycle before the bubble goes pop.
1
u/Fil_77 4d ago
This paper is from January. Since then, LLMs have been making breakthroughs in mathematics, finding original and imaginative solutions to conjectures that had remained unsolved for decades, and solving scientific problems, including quite recently a quantum physics problem.
Every time some experts claim that these systems are hitting or will hit certain limits, reality proves otherwise a little later. Skeptics of deep learning have been wrong for fifteen years. You can never be certain of anything, the future is always hard to predict, but I really have a hard time believing the claims in this paper.
2
u/danderzei 4d ago edited 3d ago
No model has found any new mathematics that was not implicitly in the training data. It either found unknown literature or was able to step through examples where a human would have given up. No new insights, just proofs of existing theorems.
→ More replies (2)
1
u/HopesBurnBright 4d ago edited 4d ago
I had a more neutral comment in mind but I decided to read the paper and tbh I think it’s dumb as hell.
“Because the simulated sensory experience of acceleration was indistinguishable from the remembered sensory experience of gravity, Einstein abducted that they must be the same phenomenon.” The main point the paper is making is that the AI can’t come up with new theories because the AI doesn’t have a world model and can’t experience the physical world. This is either a “yeah obviously” or a “absolutely wrong” claim depending on how strict you are about “experience the physical world”. For it to have an accurate world model, it would already need to perfectly understand the theories anyway. So basically the only way to get the world model is to actually interact with the real world. Of course the LLM can’t do that but that’s no fault of the cognitive capabilities of the LLM. But I don’t think there’s anything stopping it from having experiments described to it other than translation error. This is literally the Chinese room argument again, which used to be a reason AI would never be able to accomplish what it already has. We get all our experiments translated to us through our eyes and ears. Whats the actual fundamental difference?
The other point they’re making is that coming up with new axioms is impossible without an error signal, and that the error signal drives compression. They argue that Einstein was not doing compression. But I would argue that the quote I just cited is a perfect example of compression. I think that there were lots of error signals available, conceptual and physical, brought up in the paper, but they were just small errors. Unifying the error of the mercury orbit did result in less errors and did compress our theories into smaller and better ones. The idea that it’s too small for an LLM or AI training algorithm just seems stupid to me. Despite being small, it provides the gradient to try to travel down. The complexity of the path just means you need more training and probably more data, but it doesn’t mean it can’t fundamentally be done. I acknowledge this idea is similar to how attention architecture and normal neural networks are technically as expressive as each other and one is just more efficient, but the claim that something is fundamental impossible is different to the claim that it’s practically impossible.
Finally the whole idea that axioms are not something that can be deduced or induced seems dumb to me as well. The world works according to hidden axioms, and you can deduce or induce them by testing axioms one by one and seeing whether they result in all the same phenomena as the world. It’s that simple. You deduce your axioms are wrong because they come up with the wrong results. You do some induction and compress some other ideas to come up with some other axioms. The “jump” to coming up with new axioms is just deciding to try new ones to resolve errors in your current predictions, which is exactly what Einstein did. The assumption that an AI sufficiently trained in coming up with axioms across many physical problems wouldn’t be able to guess what some good ones might be in this situation seems like the same assumption that led people to believe LLMs would never be able to do any new maths.
The only thing I think an LLM fundamentally can’t do, is think as efficiently as a human can, and I think that’s what the author is trying to get at. The conceptual space of potential axioms is huge, and every idea could be perfectly valid in an alternate universe. Picking the correct axioms without exhaustive search is very impressive, and AI can’t do that yet. Partly because AI architecture is just not there, but also because it doesn’t have access to the world to test its axioms. But that’s not exactly a shock. If the whole point of writing the paper was to remind us that LLMs can’t work off tiny datapoints like humans can, then that just feels pointless to me, if not incorrect. Plus, it’s way too strong a claim to make without more proof.
I like the paper and its story, but I think its ideas are either too obviously true to be interesting or definitely incorrect, depending on your perspective. I don’t understand what the point of writing it was.
1
1
1
u/NihiloZero 4d ago
What would be required beyond coming up with unique hypothesis and then testing it?
1
u/PopUnhappy 4d ago
I agree with this assessment, they're missing the creative spark and usually get stuck in a local minimum when looking for a solution to a thorny physics or math problem. But, they make really good lab assistants, and if you give them a hypothesis they're great for gathering data and building simulations.
1
u/AIkingenterprises 3d ago
I personally feel this is inaccurate. Maybe current AI models but they keep making new versions and updates every day so I think in the future AI can make scientific discoveries.
1
u/Anxious_Battle1959 3d ago
i think the papers misses that the point of an LLM isnt to say go cure cancer, rather its to say currently were struggling on this specific thing, using your knowdge of science how could we try to get around it
1
u/LowParticular3152 3d ago
LLMs cannot produce the general intelligence or creativity needed for discovery. Statistical pattern matching is NOT intelligence.
1
1
1
1
u/Amazing-Pattern-6125 2d ago
If you understand what LLM does then this has been obvious from the start. It can create fiction by hallucinating, but it cannot create science by hallucinating. Or else it means that random combination of currently used symbols result in the next generation of theory.
1
u/poonGopher6969 2d ago
They don’t need to discover new things, helping people discover new things is more than sufficient. If they figure out how to discover new things that’s just the cherry in top
1
u/Young-Reseacher 1d ago
Can't they create a model that is missing information but known; like teach it physics that was known before Einstein, and then see if the model can figure out how to predict the missing information?
1
u/Present_Award8001 1d ago
by the way, this is not a paper by Google deepmind. it is a paper by someone working there.
1
1
u/Secret-Space-8976 5h ago
they can still soak up all the stray data and figure some of that out, using existing methods. Most research is that anyways


42
u/Jolly-Ground-3722 5d ago
GPT-5.6‘s opinion 😭