Four days after I published a piece arguing LLMs can't make the jump, OpenAI announced that an internal model called Astra had solved ten open prob...
For further actions, you may consider blocking this person and/or reporting abuse
This is such an accurate way to frame Astra’s breakthrough. AI can tackle tough, clearly defined open problems inside established formal frameworks, yet it still can’t build entirely new conceptual paradigms or come up with those foundational open questions on its own. Humans are still the irreplaceable bottleneck when it comes to framing judgment and paradigm-shifting creativity.
Appreciate it, Luna and "framing judgment" is a nice compression of the whole argument. Curious whether you think that bottleneck gets narrower over time or whether it's structural regardless of scale...
I think many smaller bottlenecks will gradually narrow, but this framing judgment feels structural. We can supercharge problem-solving, but the job of inventing entirely new conceptual playgrounds doesn’t look automatable anytime soon. 😂
Structural is probably the more testable claim since it predicts something specific: even ten more Astra-sized results shouldn't move it. Worth watching for.
My simple gut feeling (not based on any kind of knowledge or expertise, lol) says it's probably structural, and it's because current generation LLMs are "rigid" ...
What if someone would come up with a way to add "neuroplasticity" to AI/LLMs - would that be the fundamental 'game changer' they're apparently looking for ("AGI")?
At the same time, any 'fundamental' breakthrough of that kind might be incredibly risky and might open that Pandora's box which we'll then probably wish had never been opened - "be careful what you wish for" :-)
P.S. putting that more simply - if ever AI can do exactly the same things a human can do then we're screwed, it's the end - right, or not? I think that would be something we simply should not want ...
Spot on leob. Neuroplasticity as the missing piece is closer to the mark than "just add parameters."
And I'd separate the 2 claims in your P.S. parity with a human isn't the same as unsafe.
The scarier version is an agent updating its own goals faster than anyone can audit not one that simply matches us....
"Solving hard problems inside a conceptual world is not the same act as inventing the world." That distinction is so fundamental.
Having Lean 4 formally verify a non-sofic group counterexample proves that raw search within a closed, rule-bound system is scaleable. However, knowing which questions are worth asking—or recognizing when a framework itself needs to be rewritten—remains a uniquely human boundary. Loved the breakdown of Terence Tao's interaction as well; it shows how much human intuition is needed just to steer the output!
The callout regarding Lean formalization is super spot-on!
Even with a zero-sorry certificate, someone still has to audit whether the formal definition in Lean 4 truly captures the mathematical intent of the original open problem. The contrast between verifiable, closed domains (compilers, checkers, math engines) and open, ambiguous, real-world systems highlights exactly where current models excel and where human judgment remains essential. Excellent follow-up article!
The Lean-verified-but-is-the-formalization-right gap is the part worth sitting with longest — nobody's automated that judgment call yet.
The gap you point to is the real one: solving a stated problem is very different from knowing which problem to state. In my own work the model is strong at execution once the question is framed, but framing is still the part I cannot hand off. That distinction between answering and questioning is worth more attention than most benchmark scores.
"Can I hand off" is probably the sharper diagnostic than "can it answer" — worth testing against your own work: is framing unhandoffable because the model can't do it or because it hasn't been asked in a form it could act on?? I don't have a clean answer either.
"That's the line: solving hard problems inside a conceptual world is not the same act as inventing the world." - yeah I think that's spot on, that's still the difference between what the "robot" (no matter how smart) and the human can do ...
Right nd it's a good, clean way to put the whole thing — "smart" and "can invent the world" turn out to be separate axes not degrees of the same one.