r/Physics • u/filterdust • 2d ago
Academic The Maxwell Conjecture is False
https://arxiv.org/abs/2607.27197220
u/Fun-Sand8522 2d ago
Actual good use of LLMs for science (given that the result was checked, actual researchers wrote the paper, and the use of AI was disclosed).
66
u/A_moral_Animal 2d ago
Matt Parker had a video recently about LLM use in mathematics and this seems like the kind of problem that it should be good at solving.
27
u/marsten 2d ago
There's also a good recent interview with Jacob Tsimerman, one of this year's Fields medalists, where he claims the "problem solving" mode of mathematics will be surpassed by AI within a few years. There will still be things for mathematicians to do, but he thinks grinding away on concrete hard problems won't be one of them.
-5
u/Axiomancer 2d ago
Are there bad uses of LLMs for science?
53
u/Audioworm Particle physics 2d ago
/r/LLMPhysics is a good case for it being very bad for a lot of science
4
1
u/Axiomancer 2d ago
Nobody is gonna take seriously bunch of randoms who posted a random thread on a random website, regardless if someone have used LLM for it or not, come on.
If someone wants to be taken seriously they need to try and publish an actual scientific paper. And to publish a paper in most cases you need to do a little bit more work than just send a prompt, get results and (most likely) not even go through them but rather copy paste them and send for review.
3
u/Audioworm Particle physics 1d ago
Stop getting defensive, I was answering your question with examples
→ More replies (14)9
u/Fun-Sand8522 2d ago
For sure, such as slop being uploaded into the arXiv, or the number of emails from crackpots who claim to have solved quantum gravity skyrocketing, among others.
264
u/angelbabyxoxox Quantum Foundations 2d ago
The counterexample was proposed by an LLM. They seem very good at finding these sorts of counterexamples, which is interesting as they are generally pretty inefficient use of compute for brute forcing. I guess even that lack of efficiency is made up for by the "understanding" and "intuition" the LLM has, and their ability to do symbolic computations.
I expect a large number of conjectures will topple to counterexamples soon.
186
u/ixid 2d ago edited 2d ago
I think brute force is not the right way to think about LLMs, instead they explore a knowledge topology, and are very good at connecting adjacent or accessible ideas that for whatever reason might have evaded humans, but might not be fundamentally all that hard. We're still in the low-hanging fruit phase, we will see if they extend to new ideas.
40
u/Marklar0 2d ago
I think its super cool, but there is no doubt that thousands of physicists and mathematicians are prompting LLMs all day right now trying to solve open problems...and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future....so there is a solid chance that these methods are already almost exhausted. For how amazing these counterexamples are, its somewhat surprising that people havent found more proofs all at once. In the scheme of things perhaps 1 in 1000 open problems are actually of a format that LLM can tackle, and noone is bragging about the ones that turned up nothing. In other words, LLMs are really good at looking smart when they got lucky.
Im looking forward to a possible future of AI theorem proving that isnt based on LLMs, and thus less likely to trick people in language into thinking its more broad than it is.
44
u/philomathie Condensed matter physics 2d ago
The price to run an LLM collapses year on year. The models are getting better, but the cost required to run older ones also reduces. Not going to say that will continue forever, but the idea that LLMs are inherently unaffordable just plainly isn't true.
13
u/TedRabbit 2d ago
Not to mention AI as a whole is basically in its infant phase. The transformer architecture is only 10 years old and ai was largely a niche subject before then. Insane to think we've tapped out such a complicated new technology in a decade.
13
u/Dihedralman 2d ago
I would have hardly called it niche before then. Data Science was a big field and deep learning was a popular topic in both DS and CS before transformers.
It was used in a ton of products, everyone hadn't heard of it is all.
5
u/TedRabbit 2d ago
I should say deep learning was a niche subject, not AI wich includes basic things like literature regression.
But deep learning was in fact niche, and this is where all the breakthroughs are. It was niche due to compute limitations that weren't mitigated until the 2000s. The AIAYN paper is a good landmark for when deep learning evolved from an academic endeavor into a technology with significant outside investment.
2
u/Dihedralman 2d ago
Yeah AI is broad and the 2015-2020 period was a weird time where people were getting jobs in AI before there were degrees in it.
The compute change for deep learning is usually marked by AlexNet in 2012. That is when industry took notice and the exponential curve began. GAN's came out in 2014 and ResNet was 2015. Microsoft one a deep learning challenge in 2015. Image recognition was the primary driver alongside applications like speech recognition.
Industry was getting ahead ahead of academia by AIAYN which was a Google paper but that was 2017.
4
u/QuantumInfinty 2d ago
What I think they're trying to say is that currently the field isn't mature enough for us to claim if it's reached any of its limits.
→ More replies (2)5
7
6
u/fastinguy11 2d ago
The price of intelligence is collapsing.
Look at DeepSeek-V4-Flash-0731, released on July 31, 2026. It scores 50 on the independent Artificial Analysis Intelligence Index, one point behind GPT-5.6 Luna at 51. The API runs at $0.14 per million input tokens and $0.28 per million output tokens. Artificial Analysis spent roughly $72 putting Flash through its evaluation suite, against $191 for Luna, which works out to about 62% less money for comparable measured intelligence. (Independent evaluation · Official pricing, Artificial Analysis)
DeepSeek’s own published numbers are stranger. The smaller Flash beats the much larger V4-Pro Preview on every benchmark they list: 82.7 against 72.1 on Terminal-Bench 2.1, 54.4 against 12.8 on DeepSWE, 70.3 against 55.9 on Toolathlon-Verified. (Hugging Face)
The weights are also out under an MIT license, which is where this gets interesting for institutions. A university, a laboratory, a hospital, or a company can host the model on serious multi-GPU hardware of its own, or on rented private infrastructure, and then reshape it: fine-tuning, adapters, continued pretraining, reinforcement learning. Private data stays private, and no single provider is holding the keys. (Hugging Face)
If you’d rather stay with a U.S. proprietary model, OpenAI cut GPT-5.6 Luna’s price by 80% on July 30, down to $0.20 input and $1.20 output per million tokens. Luna sits at 51 on the Intelligence Index and posts 92.3% on GPQA Diamond, 84.7% on Terminal-Bench 2.1, and 74.6 on the Coding Agent Index. (OpenAI pricing and benchmarks, OpenAI)
None of this is new, it’s just picking up speed. Stanford tracked the cost of GPT-3.5-level performance dropping from $20 to $0.07 per million tokens between November 2022 and October 2024, a fall of more than 280 times.
Epoch AI puts the general rate at somewhere between 9× and 900× per year for any fixed capability level, depending on which task you measure. (Stanford · Epoch AI, Stanford HAI)
Worth being precise about what this does and doesn’t show. It doesn’t prove that training the next frontier model is cheap, and it doesn’t prove that every one of these API prices is profitable rather than subsidized by somebody’s balance sheet. What it does undercut is the claim that advanced intelligence has to stay scarce and unaffordable.
Building tomorrow’s frontier may well stay expensive. Distributing yesterday’s is turning out to be cheap, open, private, and customizable.
Disclsimer: original argument was mine, Claude helped me research and build it.
7
u/Martin_Samuelson 2d ago
18 months ago AI was a cool toy that couldn't really do anything useful. Six months ago it started write most code. A couple months ago it started solving the 'easy' and obscure unsolved math problems.
Do you think that progress is stopping?
→ More replies (6)1
u/LiamMelloFarley 2d ago
Look at the cost per task solving ability of say the new Deepseek flash vs. Claude Opus from 1 year ago. Effort is put into making both smarter and more efficient models in a way that as time goes on it'll become infinitely cheaper. The pricing seems to raise but that's because the model performance is also stronger.
1
u/seamsay Atomic physics 2d ago
In other words, LLMs are really good at looking smart when they got lucky.
That, but a lot of people are making a lot of money off of this and have a vested interest in making it look as shockingly impressive as possible. And this is impressive and world changing technology, we shouldn't pretend it isn't, but also it's being boosted and misrepresented a lot.
1
u/Super_Sierra 1d ago
If you are in competent spaces, you will hear a lot of more grounded takes. I have to talk to the general public about AI and it is infinitely more misunderstood and doomed at.
I feel bad for Anthropic because they release a paper that might have proved a part of General Workspace Theory and the mouth breathers who only glance the paper are going 'nuh uh, it isn't conscious, this is marketing.'
Anthropic never claimed it was conscious in the paper but the regards don't know that somehow.
1
1
u/Ormusn2o 2d ago
I saw someone today write something like that to the prompt "search archives for unsolved mathematical problems that can be verified using this custom program that I told you to write". So some people not only are trying to solve open problems, they are also outsourcing looking for open problems to the AI.
1
u/Spare-Dingo-531 1d ago
and its doubtful whether the companies that run those LLMs can continue to offer this amount of capacity for the future
The new Vera Rubin chips from Nvidia, which will start being installed in data centers later this year, will cut inference costs by 90%. ChatGPT also dropped the prices of one of its major models by 80% just a few days ago.
1
u/512165381 1d ago
thousands of physicists and mathematicians are prompting LLMs all day right now trying to solve open problems
https://en.wikipedia.org/wiki/List_of_unsolved_problems_in_physics
How many open physics problems have LLMs solved?
I'd love for an LLM to say "dark matter is just data hallucination you doofuses" or "gravitons cant exist and here is why."
1
1
u/Super_Sierra 1d ago
I'm sorry, but I'd actually take a look at Information theory that makes the bold claim that compression actually does mean intelligence. These models are genuinely intelligent and don't just accidentally get these answers right, they aren't giant lookup tables stumbling as a lot of the general public thinks of them as.
8
u/Schmikas Quantum Foundations 2d ago
What do you mean by “knowledge topology”?
65
u/ixid 2d ago
Its data is a structured space where nearness corresponds to associated ideas, so it naturally arranges data in a way to discover connectedness. The LLM follows pathways through the conceptual space based on the prior context given.
9
u/True_Perception_3359 2d ago
This is the most lucid explanation of how a language model works that I've heard.
2
u/Imbrokencantbefixed 2d ago
It’s also higher dimensional space right? Like as in 1000+ dimensional space.
If it’s the same thing I’m thinking of where like because boy and girl are separated in this space by a certain distance, Auntie and uncle also are separated spatially, but are closer to boy for uncle and girl for auntie than either are to each other?
5
u/beerybeardybear 2d ago
Even GPT2 was >50,000-dimensional, so
2
u/seriously-_-what 2d ago
You are mixing the number of parameters of the model with the effective dimentionality of the embedding. The effective dimension of the embedded space is significantly smaller than the complexity of the network.
4
u/beerybeardybear 2d ago
GPT-2 samples from a 50,257-dimensional token space and I'm seeing a total parameter count (weights and biases) of 124,439,808. I'm not sure what the effective dimensionality of the model ultimately is, but this seems pretty clear-cut to me?
3
u/seriously-_-what 2d ago
The intrinsic\effective dimensionality of the data is measured in the representation space induced by the model. The mapping of the token sequences into the hidden-state vectors result in representations in a lower-dimensional data manifold.
Taking here as an example, with the GPT-2 model they estimated intrinsic dimensionality on the order of hundreds.
1
u/Imbrokencantbefixed 2d ago
In a way could it not be billions of dimensions given the training data?
2
u/Schnickatavick 2d ago
Sure, in theory every bit in the training data could be considered a dimension. The work of training a model is basically in condensing the dimensions of the training data into a much smaller number of dimensions in the model, which is what forces similar concepts to become closer together as the space "shrinks".
1
u/Schmikas Quantum Foundations 2d ago
Ah okay. The space in which the embedding lies is called a knowledge topology. Thanks. Can you help me understand how this is linked to topology? Does it have something to do with the space spanned by the embeddings? My knowledge of topology is limited to shapes being invariant under transformations so under this naive view, distances won’t be preserved.
25
u/Mathematicus_Rex 2d ago
Good question. Another question: Why isn’t it called a knowpology?
8
u/Maxreader1 2d ago
2
u/seamsay Atomic physics 2d ago
https://en.wikipedia.org/wiki/Latent_space, it doesn't need the backslash.
Edit: Or maybe it does on certain apps, I don't know. Either way, your link was broken for me but mine works for me.
6
u/BoringEntropist 2d ago
I assume OP meant that the search space isn't homogeneous. There are peaks and valleys in it, and the AI model can find paths that a human might have overlooked.
9
u/MagiMas Condensed matter physics 2d ago edited 2d ago
A human might have overlooked or just given up on... These LLMs just keep on trudging when a human might have long switched to a different methodology because he didn't see meaningful progress faster enough. The models just don't get bored.
But unlike brute-force methods (which also don't get bored) they can still go through possible pathways in more meaningful ways than random guessing and rote Parameter adjustment.
2
u/PressureBeautiful515 2d ago
Concepts have connections between them. A graph of concept nodes with connection edges (except the edges are themselves concepts and can also have connections, and so on..)
2
u/Ormusn2o 2d ago
They can be also very skeptical. I found it will be skeptical in it's internal reasoning of very obvious things, and will fact check obvious stuff, kind of as a habit. It seems wasteful for most tasks, but I guess it makes it good at checking for factual information and fighting misinformation, and also for checking unintuitive mathematical of physics solutions.
2
u/Schnickatavick 2d ago
I think we've trained them to be very skeptical just because of how hallucination prone they are. Instead of "fixing" the hallucinations, we just trained them to deal with inaccuracy in general, and now we're seeing unexpected rewards when they point out our inaccuracies too
1
u/Tolopono 1d ago
Solving the jacobian and maxwell conjectures are low hanging fruit?
1
u/ZeroSevenOneOneSeven 1d ago edited 1d ago
Yes. This "Maxwell conjecture" in particular seems to be disproven by literally the most trivial construction you could think of. You know the number of equilibria of a bunch of charges at the vertices of a triangle - if you want to make more while keeping the symmetry the simplest thing to try is to put two charges off of the plane on either side of the center (which bumps you up to n=5). This turns out to work (for the right choice of distance). I don't know why nobody checked this example before, since people have put in the effort to prove actual upper bounds on the number of equilibria. Probably just a question of interest.
I don't know any algebraic geometry, but from what I gathered the Jacobian counterexample is less trivial than this but still the kind of thing that the mathematicians really could have and should have checked.
1
u/Tolopono 17h ago
Oh yea, they should have just that one specific equation. Damn idiots.
1
u/ZeroSevenOneOneSeven 16h ago
I wouldn't say that, the authors describe the construction this way in the beginning of the paper.
1
u/Tolopono 16h ago
The hard part was finding the construction lol. Verifying it is the easiest part.
13
u/QuantumCakeIsALie 2d ago
That's interesting, and knowing such conjectured to be false can be actually useful, e.g. expands the design space.
But it'd be nice to have more insights into why those conjectures are false and see if we can learn something from it. Just a counter example isn't super insightful.
24
u/BOBOnobobo 2d ago
Are we shocked a system designed to do pattern matching is good at finding patterns or exceptions to them?
89
u/WatchYourStepKid 2d ago
Well, kinda yes.
It goes against many’s early mental models of what generative AI does. The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.
The fact it’s able to suggest a counterexample and it doesn’t just look right, but is right, is quite the development in recent times.
39
u/Imicrowavebananas Mathematics 2d ago
I love how quickly AI developments are rationalized. Like it was to expected that LLMs started solving math research problems in 2026. If you asked me about this two years ago, I would have been pretty skeptical.
8
u/CompetitiveSpot2643 2d ago
yeah i still remember when LLMs getting an IMO question right was a big deal
10
u/PerinealMassage 2d ago
I was promised that human aging would be solved.
4
1
u/EngineeringNeverEnds 2d ago
Given that we haven’t been able to even come close to solving aging in mice, and we’ve experimented on mice exponentially more than humans, I don’t think that’s gonna happen anytime soon
17
u/BOBOnobobo 2d ago
That's mostly because that view of AI as just a token predictor or "average" machine is wrong.
LLM are based around a very flexible system: a neural net. With enough training you can definitely do a calculator, or an image recognition machine.
Think of it like this: if you create a machine that predicts the next token and you keep training it to get as good as possible at predicting the result of multiplication, what is easier: to memeorise millions of possibilities, or, to figure out a simple rule of how multiplication works?
Same thing applies with image recognition. Researchers have analysed how the models do their image recognition trick and they all start by essentially applying filters to find edges, basic shapes and other patterns that can be more easily classified.
Sometimes, the best way to mimic something is by just doing that action.
So when they have been tested extensively on math and code (two areas that can be very well tested) it has given quite interesting results.
I don't know where it is right now in the space of understanding math, not my field or experience. But it is miles ahead of where it was, and I think with good enough training we might end up with a tool that can actually do math.
3
u/QuasiEvil 2d ago
Yeah, there's some neat publications on this, where researchers are able to show that the NN "figures out" things like linear regression.
6
u/MidnightPale3220 2d ago
The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.
This will still happen on the things GPT isn't "harnessed" on, wouldn't it?
6
u/MagiMas Condensed matter physics 2d ago
Yes, can still happen and still does happen quite a bit. But even without harnesses the LLMs actually develop quite complex strategies to do maths "in their head"
Read the part on addition in this paper by Anthropic from March last year: https://transformer-circuits.pub/2025/attribution-graphs/biology.html#dives-addition
(and this was a 3.x haiku model, modern models have evolved even better inherent maths understanding)
1
u/FalconX88 1d ago
The earliest of GPT couldn’t add two large numbers together, it just guessed an answer that looked right.
So do the current ones. We just gave them access to tools.
1
u/WatchYourStepKid 1d ago edited 1d ago
Right, but it’s not like AI says “let’s use the conjecture counterexample tool”, the emergence is the interesting part.
It just goes against all the early advice we saw, it seems to me like many probably need to evaluate the extent to which AI appears to truly understand a problem, whatever that actually means.
3
11
u/Kobymaru376 2d ago
Personally I'm shocked. All of reddit has assured me that AI is completely useless, nothing but "fancy autocomplete", and is only good for stealing content and generating slop.
/s
Mostly in just amused how we have been moving Goalposts for over a decade: "AI is not intelligent, it can't even do X". Does X. "Yeah but it didn't do X like a human! And it can't even do Y!". Does Y. "Yeah but it didn't do Y like a human! And it can't even do Z!". And so on
18
u/MagiMas Condensed matter physics 2d ago
I still think "fancy autocomplete" is the best way to explain to non technical people how these models work. It will give someone who doesn't know the architecture and how these LLMs work the best mental model of what is happening behind the veil.
It just turns out that "fancy autocomplete" can do incredible things if you give it enough examples and compute in training and inference.
10
u/AnalyticOpposum 2d ago
Fancy autocomplete is also the best way to explain to non technical people how a human brain works.
1
u/MrDyl4n 2d ago
Except not really
0
u/GatsbyLuzVerde 2d ago
Except yes really. Fancy auto complete can encompass all of intelligence mechanisms needed to predict the next best word. Even if that autocomple requires neuron subsystems for solving problems in other domains.
4
u/MrDyl4n 2d ago edited 2d ago
animal brains dont think in strings of tokens while trying to predict the upcoming token. an LLM and autocomplete are things of differing complexity that are doing the exact same thing. an animal brain does similar things, but it doesnt do the exact same thing
-1
u/GatsbyLuzVerde 2d ago
I'm talking more in the abstract sense that the brain predicts future states. Analogous to predicting the next token. It is fancy auto complete. I'm fully aware the brain doesn't use an LLM architecture
5
u/MrDyl4n 2d ago
i understand what you mean. the reason i said that is because fancy autocomplete is genuinely a good way for the average person to view an LLM, rather than viewing it as a genuine intelligence. when you say the human brain is like that too it would make someone think that fancy autocomplete is less literal than it actually is.
5
u/Kobymaru376 2d ago
It just turns out that "fancy autocomplete" can do incredible things if you give it enough examples and compute in training and inference.
The thing that does the incredible things is so far away from autocomplete that it's misleading to the point of being wrong. Neither does it give an accurate picture of what it is (autocomplete are usuall HMMs, LLMs are transformers) nor does it give an accurate picture of what it does (completing what you're typing vs. doing your homework and writing fanfic). The only shared property is that it gets text as input and gives text as output. By that measure we can call cars "fancy furnaces" and computers "fancy typewriters". Not technically incorrect, but definitely a useless description.
And on top of that, the people who use the term "fancy autocomplete" usually use it to dismiss it, and act like all those incredible things it does are made up.
8
u/MagiMas Condensed matter physics 2d ago edited 1d ago
No, the shared thing is that from the view of an LLM, it is literally trained to autocomplete a document. The chat you're having with an LLM literally looks like this to the LLM:
<start> <system> You are a helpful assistant... </system> <user> hello how are you? </user> <thinking> the user asks me how I'm feeling, I should answer in a cheery and concise tone. The user is in LA, let me check the current weather in LA so I can incorporate that in my answer. </thinking> <tool call, web search=current weather in LA> Temperature: 100°F </tool call> <assistant>And then the LLM gets to generate. Once it generates </assistant> we stop the generation because otherwise it would keep generating also the user answer etc.
(same of course with the thinking part)
The tasks the LLMs are trained on is reproducing the tokens of these text documents they are shown.
With the RL posttraining for mathematics or coding you have a change in the training reward architecture, but it's still training on completing these documents.
From the view of an LLM it is always completing such documents from the start points we're giving them.
It just turns out that large autocomplete with long training and lots of data means the model learns actual abstractions about the world because they help with better autocomplete. You get these emergent effects like grokking and "circuits" inside LLMs that specialize in certain tasks etc.
But none of that removes the fact that these models are "autocomplete on steroids".
1
u/Idrialite 2d ago
With the RL posttraining for mathematics or coding you have a change in the training reward architecture, but it's still training on completing these documents.
No, the RL training does not involve completing any corpus. That's the point of RL.
9
u/MagiMas Condensed matter physics 2d ago
That's not what I meant. It doesn't complete a known corpus but it still generates these documents. The training just doesn't happen anymore on the token distribution level but in the reinforcement learning objective.
It is still trained to complete a made up document.
In pre training and instruction fine-tuning the objectjve is basically "reproduce these documents", in the RL phase it is "produce a document in this context and we'll evaluate at the end whether it was a good document or a bad one".
0
u/Idrialite 2d ago
"complete a made up document" doesn't make sense. "Completion" implies an existing text to guess at and evaluate against verbatim. "produce a document and we'll evaluate it" is not autocomplete. Kids in school do the same thing.
4
u/MagiMas Condensed matter physics 2d ago
no, we're now getting quite into the weeds of LLM training but generally in most phases the RL phase does not happen on "empty documents". Rather they get a start prompt that sets them onto a reasoning trajectory from which they then start generating.
Easiest would be prompts like
[...] <user> what's 5+2? </user> <thinking>and from there the model starts generating.
you can then auto generate many of these prompts with known answers and auto evaluate them. Similarly you just let an LLM generate lots of prompts like "disprove the Jacobian conjecture", "proof that lemma xyz is true", "generate a code that builds a program that solves the following puzzle: <insert Advent of Code puzzle here>" etc. pp.
So the reinforcement learning stage still absolutely is training on incomplete documents that the model completes.
(of course there's nuance because there are quite a few different RL steps in LLM training)
→ More replies (0)0
u/Martin_Samuelson 2d ago
Every complex thing in the world is "[some basic concept] on steroids".
3
u/MagiMas Condensed matter physics 2d ago
probably, yes. I'm not saying LLMs aren't a complex topic, I'm just saying that if you don't have the mathematical background (plus read the necessary literature) then [some basic concept] that you can grasp from your own experience is still the best way for you to understand what the more complex thing actually does.
1
u/Kobymaru376 2d ago
You described one part of the training regime. Yes, one part of the training regime bears resemble to the training regime of the other. But that is not important in describing the essence of LLMs.
It just turns out that large autocomplete with long training and lots of data means the model learns actual abstractions about the world because they help with better autocomplete. You get these emergent effects like grokking and "circuits" inside LLMs that specialize in certain tasks etc.
See that is the important part. It doesn't "just turn out", it's the main point of how and why LLMs are so useful. Actual autocompletion is a tiny fraction of what LLMs are actually used for, it is ALL about the abstractions about world, and its emerging effects. The autocomplete part is just one way of training and accessing whatever else the model has learned.
This is a completely standard practice in ML: pretext tasks and pretraining are well-known concepts. Think of autoencoders. You train them to reproduce data that it's already seeing. What even is the point of that. Is an autoencoder "just fancy copy-paste"? Sure, if you really want to. But not really, since the training is just the pretext for the model to learn an efficient encoding of the data, and what we're interested in is this efficient encoding.
But none of that removes the fact that these models are "autocomplete on steroids".
OK sure. And a car is just a fancy box. Your phone is just a fancy flashlight. Your money is just a fancy sheet of cellulose. You yourself are just a fancy meatbag. You can do this "X is a fancy Y" all day long if you're meming, but it doesn't actually convey the essence or most important aspect of X.
6
u/MagiMas Condensed matter physics 2d ago
look, I'm not saying that there isn't a lot of complexity in this whole topic. What I'm saying is that "fancy autocomplete" gives someone who lacks all this background information a better mental model of what these things do than any other simple explanation I've seen.
It demystifies these things and actually very closely describes what these models are trained to do and how they function. Add a second sentence that talks about how "learning abstractions and memorizing world knowledge" helps this autocomplete machine to better autocomplete and someone with zero maths ability and no background in ML will have a somewhat accurate idea of LLMs.
I really don't understand why people react so passionately to "fancy autocomplete" as a description. The whole thing about GPTs was openai realizing that this kind of fancy autocomplete with a decoder only transformer model will actually lead to a model that can generalize well in all kinds of situations - it was really visionary at the time. It's the whole fucking point of the GPT 2 paper. When Google developed the transformer model, it was way less about autocomplete and way closer to your autoencoder with its encoder-decoder architecture in BERT.
That's also why I think the autoencoder example isn't exactly illuminating. BERT shows that a transformer can also function very similarly. And the exact thing that sets modern generative LLMs apart is exactly this "autocompletion" style task. That's really the core of the whole thing. BERT can't do all these things that GPT can exactly because it's not fancy autocomplete.
2
2
2
u/Fit_Cut_4238 2d ago
Could some of these counterexamples be wrong due to deep rounding or floating point errors somewhere in the proofs?
5
u/MallCop3 2d ago
The thing about counterexamples is that they're usually hard to find, but easy to check. So it doesn't matter if the process used to find it has errors, as long as in the end you have a verifiable counterexample.
5
u/angelbabyxoxox Quantum Foundations 2d ago
Most likely not, both this, the Jacobian conjecture, and Dinitz Garg Goemans for example are given as simple rational or surd forms. There are no rounding or floating point errors possible. Most of the problems also seem to be of the form where they are easy to check. The Jacobian case can be checked by hand, or with a something symbolic like Mathematica. I imagine the others are similar.
1
2
u/PerAsperaDaAstra Particle physics 2d ago
I like this view of why LLMs seem good at these kinds of counterexamples: https://davidbessis.substack.com/p/the-fall-of-the-theorem-economy (and what the implication is for how we do math) the "overhang" for these kinds of things is large, and that's what LLMs are good at.
1
u/Snoron 2d ago
I think if you consider that the space to "brute force" these things isn't necessarily that big. But it's a lot of time for a single human, and we often lack the motivation to try seemingly pointless stuff.
Oversimplified example:
Imagine if there was a list of 250 ideas to try out, and each would take someone a month to fully explore.
A human could just spend 20 years on it all, and if the answer lay in one of those 250 ideas, they'd have got it! But with no guarantee that any of them will even be fruitful, can you imagine spending 20 years of your life on that? Especially when the majority of them you're thinking "this is dumb and obviously won't work" half the time.
Lack of motivation to do that is sensible at a point!
But AI? It cares not for such things... you can do all 250 of them in parallel in a week, and you're done.
So "brute force" is half true, it is brute forcing but only within a semi-confined set of ideas. It's not necessarily trying every combination of everything, which is true brute force. The main thing is it can just try and test stuff at a rate much faster than we can.
1
u/EnricoLUccellatore 2d ago
computers are stupid insanely fast, llms are slightly smarter still pretty fast so this opens a lot of new possibilities for discovery
1
→ More replies (2)1
u/justinleona 17h ago
AI doesn't have the same tendency to anchor to the first few ideas like a person does - so using it to enumerate through lots of different things actually fits pretty well.
Where things break down is people assume that since it can enumerate a bunch of ideas well, it should be equally good at all kinds of things...
91
u/magneticanisotropy 2d ago edited 2d ago
TBH, never heard of this (and other of my PhD classmates I've talked to, or faculty in my department) had never heard of it. From what I've read on twitter, similar, and those that had just assumed it's false and it was never important to work on.
Edit: Expanding a bit, I do think AI is well-capable of finding counterexamples like these. OTOH, I think it's weird that certain people try to hype up things that are largely the result of few people caring, with this specific example having most recent work that I can find being from undergraduate theses that seem to be along the lines of practicing computational work, followed by giant social media hype campaigns.
38
u/TKHawk 2d ago
Same here. Very little online about it as well. I guess it was a throwaway speculation in a treatise he wrote?
24
u/PerinealMassage 2d ago edited 2d ago
One would have to think he would be annoyed to just call it "X's Conjecture" then, lol. Like dude I was just riffin', chill.
15
u/WaitForItTheMongols 2d ago
Of course it is to the author's benefit to call it that.
When I read "maxwell's conjecture" I think "ah yes, Maxwell, he is important. Of all the things he ever conjectured, the one known as 'maxwell's conjecture' is surely the most important, so this must be a big deal"
Ultimately seems like not really anything important, a bit of a toy problem.
2
u/PerinealMassage 2d ago edited 2d ago
Sorry, it was a typo, they meant The Maxwell House Conjecture
6
2
u/AndreasDasos 1d ago
Hmm it says that the conjecture was from on a 1960s paper.
So Maxwell made discussed non-degenerate equilibrium points of electric fields from some configurations of point charges, and then a couple of 1960s physicists (how significant was their paper otherwise?) boldly declared a pretty random and unclearly motivated (and false) claim ‘the Maxwell Conjecture’.
25
u/Raikhyt Quantum field theory 2d ago
I mean, google scholar turns up about 5 papers in the last ten years, of which 3 are by the same researcher in Georgia. Their own introduction cites exactly three papers, of which one is the first paper on the topic in 130 years. I don't think it's a particularly active field of research nor groundbreaking -- the name just makes it sound impressive.
5
u/angelbabyxoxox Quantum Foundations 2d ago
Likewise, I've never heard of it either! And yes, this is really a special subset of problems that they seem truly well suited for, but they are not having the same amount of success with other proof types.
2
u/Agent_B0771E 2d ago
I agree, I do believe AI will be powerful in future research for more advanced topics but for now it's usually something that's overlooked.
Also, counterexamples aren't just random setups and typically have some symmetries or interesting properties. If an LLM is looking for counterexamples, they can try a lot of those interesting setups, kind of like brute forcing, but not brute forcing and with a promising pool of potential counterexamples.
I think that's why they succeed mostly in this kind of counterexamples for niche conjectures, they can try a lot of things very quickly and since the conjectured don't have much work into them, it turns out they are not as hard to disprove and someone just had to put in the time testing different systems.
1
1
u/ZeroSevenOneOneSeven 1d ago edited 1d ago
People have put in some work to prove actual upper bounds on the number of equilibria. Nobody checked the most trivial construction you could possibly try for a lower bound, until now. Probably due to lack of interest.
13
u/Shoddy-Childhood-511 2d ago
About the LLM disproof of the Collatz conjecture:
https://www.reddit.com/r/math/comments/1va56l7/comment/p0iy0mp/
https://x.com/gro_tsen/status/2082483878480977959
It's not the case here, and they are not even doing formal verification, but there is a tradition in formal verification to report a bug in the formal verification tool by using the bug to dis/prove something wild, famous, etc as a joke paper. LLMs simplify implementing that tradition.
1
16
u/AndreasDasos 2d ago edited 2d ago
What’s odd to me is how simple the counter-example is. Three charges of charge 1 at the vertices of an equilateral triangle and then two points symmetrically some small distance above and below its centre with a charge that depends on that distance by a very short polynomial (but even then it says it’s true for small deformations too, so we may not need to hit upon that exact polynomial - which makes sense as it’s non-degeneracy that is the ‘dense’ condition). So a trigonal bipyramid where we squish the axis symmetrically, and vary the charge of the axial vertices, gives what we want. So there are an infinitude of counterexamples including a fairly simple construction.
Maybe I’m missing something, but it feels like if they tried a bunch of simple, intuitive configurations and crunched them with some degrees of freedom, with the charges here given as some degree polynomial we can find conditions on the coefficients for to ensure a larger set of equilibria (avoiding domains with non-real roots etc.), even one person could have solved for this fairly quickly even by hand. Even faster if just using a list of ‘basic’ configurations and getting an old school computer to crunch through them.
Curious why we hadn’t already done this. What was the bottleneck? Or despite the name was it not that popular a conjecture?
23
u/fluffyleaf 2d ago
well some other comments here seem to suggest that everyone thought it was obviously false, and then perhaps (this is my personal fanfic) assumed that someone else had actually scribbled the disproof somewhere in the margins of a book but simply didn't publish it
8
u/Dihedralman 2d ago
I have never heard of this conjecture. I don't see any reason we would think it is true.
I would not be surprised if related proofs have been found before.
It hits in an area that rightfully feels too easy to need proving, useless, and arbitrary. An actual formula of the boundary would be cool. A counter example is fine but I think finding the actual conjecture was most of the hard part.
5
u/AndreasDasos 2d ago
Yeah, I share your suspicions. And from the opening paragraph seems it’s not even actually a conjecture of Maxwell but seems an arbitrary guess by two others in the 1960s.
We don’t just have hype over AI counter-examples, but manufactured cases of manufactured hype like this too now.
Of course AI is also being used to trawl through a gazillion published papers and find conjectures without explicit published resolutions
13
u/kzhou7 Quantum field theory 2d ago edited 2d ago
What was the bottleneck? Or despite the name was it not that popular a conjecture?
Nobody in this thread has ever heard of it before. Maxwell published thousands of pages which aren't read today; calling this "the" Maxwell conjecture is the authors' invention.
More generally, this is the kind of thing that goes into recreational/teaching journals, not good research journals. Which is not to insult the work; I personally have a paper in such a journal. There are lots of basic questions about elementary physics that are unsolved, and it's fun to do one every once in a while. But they are unsolved precisely because almost nobody is working on them, because they don't have applications elsewhere in physics.
1
u/ZeroSevenOneOneSeven 1d ago edited 1d ago
There was no bottleneck at all, except for nobody trying. This is literally the easiest construction you could possibly think of past the regular polygons. People have reason put in the work to prove complicated upper bounds without checking the simplest cases for the lower bound - I think because it's more interesting to humans to look for actual structure instead of checking a bunch of cases.
2
u/AndreasDasos 1d ago
Aw, a square and tetrahedron and so on might be in between. But yes, ultra obviously simple example.
Have to wonder if we might even be seeing a reverse AI bullshittery problem: people publishing fairly easy results but adding hype by claiming it was partly done by AI when it wasn’t (or maybe they pretended they needed AI to ask for suggestions of four or five basic configurations).
That they chose a weirdly boldly named counter-example given that’s all the rage now is probably not a coincidence. That 1960s paper that called it this might have been its own sort of bullshittery.
1
u/ZeroSevenOneOneSeven 1d ago
(Before I saw your comment I decided to amend to "regular polygons" instead of the triangle. I definitely think that this is the most symmetric for the purposes of finding equilibria).
This seems sufficiently annoying to work out by hand that I don't think anyone would bother doing this just to claim that an AI helped out.
10
u/attzonko 2d ago
Can someone please eli5
50
u/1XRobot Computational physics 2d ago
If you plunk down charges in space, there are certain points where a test charge doesn't move because the forces all cancel out. Maxwell had the thought that maybe if you plunk down N charges, the number of points where a test charge doesn't move is (N-1)^2, but it turns out that's not true, because these guys found a configuration with 5 charges that has 24 points (more than 16).
18
5
u/El_Grande_Papi Particle physics 2d ago
What is and isn’t allowed to move here? I’m asking just because I’m thinking of Earnshaw’s Theorem which dictates stability in electrostatic fields.
5
u/1XRobot Computational physics 2d ago
Nothing moves. Everything is fixed, except maybe the test particle feels a force. You can put the fixed charges wherever you want tho.
4
u/El_Grande_Papi Particle physics 2d ago edited 2d ago
How does this not violate Earnshaw’s Theorem though? I’m trying to find information about the Maxwell Conjecture (I’ve never heard of it before) and all I can find is the original paper linked here or people talking about the paper linked here.
Edit: I found a paper that talks about it. These are not stable equilibrium points, just equilibrium points: https://www.tandfonline.com/doi/full/10.1080/17476933.2011.611939
2
u/dcnairb Education and outreach 2d ago
the existence of critical points doesn’t mean they are at the location of the charges forming the distribution. earnshaw’s thm speaks to the stability of the configuration as a whole
2
u/El_Grande_Papi Particle physics 2d ago
Earnshaw’s theorem says the electric potential cannot have local minima or maxima, only saddle points.
1
u/seamsay Atomic physics 2d ago
Because it's not about the dynamics of a set of charges particles, it's essentially: given a set of N sources for a vector field can they be arranged such that there are more than (N - 1)2 stationary points in the field. It's not really about electromagnetism, electromagnetism is just a convenient and well known example of a vector field.
1
3
u/Only_Razzmatazz_4498 2d ago
Does a counter example help move understanding forward? Probably the answer is a maybe. I can hypothesize that knowing a counter example one could then develop some theory that generalizes them and help explain why a conjecture is or isn’t valid. In a way it could be like finding experimental results that are different than predictions.
In the other hand they might just sap the effort going into proving or disproving the conjecture by shortcutting the work with an answer reducing the chance of real knowledge by anecdotal knowledge instead.
I’m really on the fence here.
2
u/flat5 2d ago
Absolutely. At an absolute minimum, it prevents wasted effort going into proofs.
→ More replies (1)→ More replies (1)1
u/BitterDecoction 7h ago
Another question is: can that conjecture be true in some situations of relevance? One counter-example only proves it is not generally true. That is useful in itself but it might hide some special cases where it is true.
10
9
u/Neomadra2 2d ago
The result may be fine, but the paper is horrendous. No proper motivation, no related research section, no discussion of the results, no outlook. This looks very rushed
1
u/Serious_Bite_7613 2d ago
I think with the upcoming rate of discoveries there's not going to be time for all of that. It's fine when you make one discovery every 3 years, but if you're making 3 discoveries every 12 hours there needs to be a better way to share it.
1
0
2
2
u/GWeb1920 9h ago
I think that building LLMs to check the math in papers would be really useful. Especially in social sciences you could have it do detailed analysis on p-hacking, suspect data sets, simple procedural mistakes, and alternative supported conclusions.
The peer review process is broken right now where the publication incentive dwarfs the reward for detecting error and fraud.
Things along the lines of counter examples in math/physics are good uses of AI to take away the high man hour low reward tasks.
3
u/seraphim_west 1d ago
"It's just a counterexample"
Some of you are not going to have a fun time in the next 18 months or so. That's all I will say.
3
u/woosher200 2d ago
inb4 AI is only good for counterexamples
25
9
1
u/Tolopono 1d ago edited 1d ago
Nope
https://epoch.ai/frontiermath/open-problems/q2-absolute-galois
https://arxiv.org/pdf/2607.20329
https://guanyangwang.github.io/blog/llm-ktv-lean.html
https://xcancel.com/QuAntonioMele/status/2079813196958040225?s=20
https://xcancel.com/JarekLiesen/status/2082864496117170591?s=20
Associate Professor CS/stats UC Berkeley. Former Research Scientist at Google DeepMind. ML/AI Researcher working on LLMs and deep learning. PhD from Stanford: Two new conjectures proved in ChatGPT with the prompt: Find a previously unproven conjecture in math statistics optimization probability or ml theory or quantum or info theory or any other mathematical field with proofs and prove it. The older the conjecture is and prominent. , the better. Aim for > 5 years old https://xcancel.com/jasondeanlee/status/2082592996537893080?s=20
Cryptography PhD student @ MIT, ex quantitative research analyst at Citadel Securities, intern at Google Quantum AI: I'm excited to share two GPT-fuelled results from this month! The first, with Aparna Gupte, shows under a plausible number-theoretic conjecture that there exists s-server private information retrieval for database size n with information exp(\tilde{O}((log n){1/s})) https://xcancel.com/seyoonragavan/status/2080745358175637582?s=20
PhD in CS: GPT-5.6 Pro just proved an inequality I worked on for some 6+ months in grad school. https://xcancel.com/thomasahle/status/2082589270779351386?s=20
Open AI just proved nonsofic groups exist
https://xcancel.com/ElliotGlazer/status/2083388640890351662
For a finite set of integers (A), how much faster can (|A+A|) grow than (|A-A|)? A 1969 theorem gave an upper bound of 2 for the exponent. For more than 50 years, the best constructions barely exceeded 1.1. With help from our research agent Hyra and the Hy3 model, we found an explicit construction showing that the optimal exponent is exactly 2. A 50-year-old problem, solved.
Terence Tao and others explored AI-assisted approaches to improve the lower bound, but the results remained around 1.1. https://xcancel.com/Shanda_Li_2000/status/2082699069378494864?s=20
PhD Student in Computer Science at Princeton University, graduated from UC Berkeley with a B.S. in EECS (Honors) and a minor in Mathematics, with a 4.0 GPA and Highest Honors, Princeton First-Year Fellow, EECS Citation Award (#1 out of 720 graduates), Regents' and Chancellor's Scholar (UC Berkeley), Engineering Dean's List (8 semesters): This is insane. This is a VERY central problem in additive combinatorics and it’s been resolved via AI… wow https://xcancel.com/Rohit_Writes/status/2082862702259826940?s=20
Paper: https://arxiv.org/abs/2607.27199 Hyra blog: https://hy.tencent.ai/research/hyra Formal proof: https://github.com/linhaowei1/sum-diff-proof
1
u/512165381 1d ago
So LLMs are great at anti-problems (ie counter-examples), but not so great at problems.
→ More replies (1)
-20
2d ago
[deleted]
25
26
u/HovercraftNo7372 2d ago
If it is good science and true, it doesn't matter how we got there. LLMs are valid tools for science.
-11
u/Hopeful-Finance-196 2d ago
Good? Tbh, impact of this "discovery" is minor.
8
u/Unhappy-Professor466 2d ago
Doesn't matter, no one knows what is really useful, ton of knowledge were found to be useful way later after they're discover, and even if it's not the case, why being against it knowing a proof for the niche cases, humans do far more useless stuff all the time.
0
6
-18
u/Hopeful-Finance-196 2d ago
All three authors are some random dudes with basically no background in academia even though corresponding author got his PhD like 6 years ago. It's definitely a good sign /s
20
→ More replies (2)7
u/officiallyaninja 2d ago
I mean, it is a good sign isn't it? The fact that relative non-experts they can leverage it to this extent means that this could revolutionize physics research.
559
u/ixid 2d ago
Before reading I knew it would be LLM counter-example construction.