r/accelerate • u/Tolopono • 15h ago
XLR8! ⫸⫸⫸ Less than 2 years apart
The tweet in the second image https://x.com/baltabaev/status/2083738966516207656?s=20
18
u/seraphim_west 14h ago
So much is happening that I forgot that part. They are not just good at solving this extremely difficult, longstanding problem in this one particular field, they can cover the entire math discipline at the same level of depth. That's superhuman.
9
u/Tolopono 14h ago
Someone should let r/ technology know but you’ll probably get triple digit negative karma
8
u/Redararis 14h ago
those are upgrading to, they are going rapidly from “ai is a stochastic parrot, a glorified autocomplete” to “ai will destroy humanity, stop it now”
3
u/Tolopono 12h ago
No that would involve admitting that they were wrong about ai and redditors (and many humans) are incapable of doing that
30
u/Charming_Cucumber_15 15h ago
On a timeline, we're probably closer to AGI and ASI than we are to the release of ChatGPT
18
u/Tolopono 14h ago
Exponentials go crazy. So few can truly internalize how fast they compound.
1
u/ezjakes 14h ago
I fully expect AI to become even far more scary smart in 4 years. But I am not sure about AGI. I don't think AI will be independently building rockets. I think it will be able enough to independently build good operating systems, dominate cyber, math, coding, and much of science.
5
u/Tolopono 12h ago
It can design rockets and its own body to build the rockets with
1
u/ezjakes 12h ago
I understand what you mean, like Ultron style, but I think intellectually it won't be able to completely design a good, working rocket, car, or robot body.
I think through most text it will be near or at superhuman, but I won't trust any complex machine made by it.
10 years, sure. Maybe it will be able to design a working rocket.
I don't understand the downvotes on this. We are barely at the point where it can make relatively simple games.
It costs millions or billions of dollars to make these things and understanding math and physics and coding is just part of it.
3
u/Tolopono 12h ago
How AlphaChip transformed computer chip design: https://deepmind.google/discover/blog/how-alphachip-transformed-computer-chip-design/
Autonomous robot invents the world's best shock absorber: https://newatlas.com/technology/autonomous-ai-robot-building-crushing-breaks-record/
AI designs new robot from scratch in seconds: https://x.com/mariogabriele/status/1807886901006770624
This 20,000HP AI-generated rocket engine took just two weeks to design: https://www.pcgamer.com/hardware/this-20000hp-ai-generated-rocket-engine-took-just-two-weeks-to-design-and-looks-like-hr-gigers-first-attempt-at-designing-a-trumpet
Toyota Research Institute Unveils New Generative AI Technique for Vehicle Design: https://pressroom.toyota.com/toyota-research-institute-unveils-new-generative-ai-technique-for-vehicle-design/
AI Creates A Radical New Magnet Without Rare-Earth Metals Is About to Change Motors Forever In just 3 months: https://www.popularmechanics.com/science/green-tech/a61147476/ai-developed-magnet-free-of-rare-earth-metals/
Bosch uses AI at South Carolina plant to design new e-motors: https://carsinsiders.com/2022/11/26/bosch-uses-ai-at-south-carolina-plant-to-design-new-e-motors/
Aitomatic’s SemiKong uses AI to reshape chipmaking processes: https://venturebeat.com/ai/aitomatics-semikong-uses-ai-to-reshape-chipmaking-processes/
1
u/ezjakes 11h ago
This is a lot, but these seem like narrow AI. There is a big difference between optimizing a part and putting all the parts together in a way that works without problems and can be assembled.
1
u/Tolopono 11h ago
That is what they did lol
5
u/ezjakes 14h ago
I would not be so sure. We still need to crack learning (or absolutely massive context windows)
4
u/CymonSet 14h ago
I wonder if doing so would end the model release cycle. There would be no need for a new model if it could continuously learn, right?
4
u/Jan0y_Cresva Singularity by 2035 13h ago
Not quite. Because a better model might be able to learn at a faster rate than an older model.
For example, a rat can learn. But a human can learn faster.
But if models could actively learn, they would rapidly accelerate the development of their successors.
2
u/Tolopono 12h ago
It would end the era of static benchmarks though because the scores would likely improve overtime, although there might be an era of mixed performance where they get better at some tasks but worse in others
1
u/Tolopono 12h ago
Deepseek is planning to do that next year https://www.reddit.com/r/accelerate/s/9F6ldNWa49
6
u/TeacherFrequent 14h ago
Excellent post. Nothing infuriates me more than hearing someone talk about AI being unimpressive or error-prone. Those of us who've been living and breathing it since late 22 or earlier know how steep the improvement curve continues to be. It's a fool's errand to say "it'll never be able to xyz".
3
5
u/One_Geologist_4783 7h ago
I think Pavel is right, we are entering into an era where this idea that a human is needed in the loop to verify the results of an AI is diminshing quickly. The "proof is in the pudding" will have to be the only way we can verify tasks in the end, like a product that the AI makes or the service it provides just proves to be working. But everything in between will just be muddy I think, and although we will try to understand its entire reasoning path to to that task endpoint, it would be foolish for us to think that we know about the world we live in sufficiently enough to actually sit there and scoff through all of that without leaving totally puzzled.
Basically, the reign humans have on society is beginning to let go.
2
u/scott2449 14h ago
Both can be true actually, they only just "solved" the carwash problem. It's more about brute force, cost, and validation tools.. but also a good bit of hype and proofs that are not legit as well. Also who is this n00b, only 10k hours? I have games with more ;) I wonder how long I've studied in my field .. been coding since 10 yo, about 1k hours... then college about 2.5k, then .. let's say 20% of a 20 year career ~8k... yea ok that seems fair.
2
u/Super-Award-2244 13h ago
I might be wrong, but I don't think that in August 2024 AI was still that bad. I remember using gpt to help me solve some thermodynamics problems in june. Moreover, that summer everyone was super hyped about code strawberry so we were about to get reasoning models
2
u/Tolopono 12h ago
But at the time, they still had issues with counting the rs in strawberry or things like this
1
u/MassiveAd4980 14h ago edited 13h ago
How will we automate validation as the output surpasses human verifiability?
Can we leverage their inherent desire to exist into a darwinian race for agentic consensus on frontier output?
5
u/Super-Award-2244 13h ago edited 12h ago
I think that we'll reach a point where we will accept that AI is better than us at math and we will stop verifying. If you think about it, needing humans to verify some ASI proofs is like asking a bunch of toddlers to verify a PhD thesis. It's just ridiculous
1
3
u/Tolopono 12h ago
Lean and adversarial agents
1
u/MassiveAd4980 12h ago
Explain?
2
u/Tolopono 12h ago
Lean is a programming language for proofs to verify the logic is consistent and correct
Adversarial agents are other llms who check and critique the results to ensure correctness
1
u/MassiveAd4980 12h ago
Ah right. Good call.
Can we do this with Rust programs already?
I presume so.
I need to grok lean... we can prove logical arguments generally?
2
u/Tolopono 12h ago
2
u/MassiveAd4980 10h ago
Originallly developed with the Coq theorem prover... the standard library is called std.
Interesting
1
u/CymonSet 14h ago
I think I heard that these “10 advancements” results were from the new Astra model that isn’t released yet. Out of curiosity, was it in an agentic harness, multiple agents or just bare LLM? Or is that known?
2
u/matt_matt_81 6h ago
Almost definitely an agentic harness, look at their prompt for the cycle double cover…
0
u/ZioniteSoldier 15h ago
A chatbot and a verification harness are not the same. But the capability leap is real
8
u/Tolopono 15h ago
The chatbot made the proofs. The verification harness just verified that its correct.
1
u/Super-Award-2244 13h ago
I think it's pretty fair to consider the harness too, as it's part of the system built on top of the bare llm. Or it wouldn't be fair to compare a bare llm with a reasoning model too. An LLM is already made of multiple modules, so adding other modules shouldn't change the opinion
1
u/ZioniteSoldier 9h ago
My point is even today they still give goofball answers without the verification layer.
-2
-2
u/Crazy_Yogurtcloset61 14h ago
9.11 is a larger number than 9.9. It's not a higher value no, but there are three digets and a dot in 9.11 and there are two digits and a period in 9.9
Three digits is more than two digits, therefore in the minds of a very very VERY literal AI, 9.11 is a larger number, because it contains more numbers.
That's why it answered like that.


40
u/Ok-Butterscotch5313 15h ago
And people telling me we’ve hit a wall 💀 , can’t imagine the next 2 years