OpenAI’s amazing — but vastly oversold — new model Astra

Eight or nine misconceptions about Astra. See if you can spot the biggest fallacy.
OpenAI’s amazing — but vastly oversold — new model Astra

Astra, a new model that OpenAI is testing internally, is amazing. No denying that:

@OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.\n\nWe believe it will be a major step for scientific reasoning.openai.com/index/ten-adva…“,“username”:“polynoamial”,“name”:“Noam Brown”,“profile_image_url”:“https://pbs.substack.com/profile_images/1872883170292457472/8ywVGO5M_normal.jpg”,“date”:“2026-08-01T08:17:42.000Z”,“photos”:\[{"img\_url":"https://pbs.substack.com/media/HOn2XyYasAA\_rAp.jpg","link\_url":"https://t.co/jHuulDwV46"}\],“quoted_tweet”:{“full_text”:“10 proofs from our next major model Astra on long-standing open problems in mathematics and theoretical computer science (also including new circuit lower bounds for computing the permanent!)\n\nGPT-5.6 has already enabled so much exciting work in math and science. Can’t wait to”,“username”:“wjmzbmr1”,“name”:“Lijie Chen”,“profile_image_url”:“https://pbs.substack.com/profile_images/1810196750009094144/nq1dEwmM_normal.jpg”},“reply_count”:664,“retweet_count”:2107,“like_count”:14941,“impression_count”:8395185,“expanded_url”:null,“video_url”:null,“video_preview_media_key”:null,“belowTheFold”:false}“ data-component-name=“Twitter2ToDOM”>

Part I: The Biggest Fallacy

But at the same time, a whole raft of people, some fairly prominent, are running around making a deeply flawed argument about the implications.

See if you can spot the fallacy. Here are three examples among many.

Another tweet went so far as to claim that “the species just crossed a one-way threshold”; Elon Musk took itas evidence that we had reached The Singularity.

Each is making essentially the same error.

§

The fallacy has a name; it’s called thefallacy of composition.

What’s the manifestation of the fallacy in the current case? Thinking that a system that is great at a certain kind of math problem is great at all math, great at science or even quite possibly great at everything. The “AGI-is-near” community keeps committing the same logical fallacy over and over. Every time there’s an advance, I see the same error.

Here’s how the fallacy works.

1. Someone pretends that all cognition is created equally. (Totally untrue.)

2. Whenever AI achieves success on some form of (fancy) cognition, they want you to believe that success on all forms of AI is imminent.

You don’t have to be a cognitive psychologist to realize that this inference just doesn’t follow. We all know, for example, that expertise in math doesn’t guarantee genius in all domains. Someone who is great at math or physics or programming may struggle with writing or understanding human relationships (and conversely a great writer may be weak at math, etc).Expertise in one domain does not at all guarantee expertise in all or even most domains.

That’s precisely *why* people like Howard Gardner and Robert Sternberg developed multidimensional theories of intelligence, why the SAT tests math separately from verbal, etc.

Astra appears to be —we still haven’t seen the methodology—great at math, or at least some forms of math, but that does not mean that it will avoid hallucinations or solve the reliability problems other GenAI systems have. It doesn’t even mean it will be able to read PDFs reliably. And it doesn’t mean it will be the first generative AI to be able to obey hard rules, either. (Which should terrify you.)

In particular, Astra is obviously excellent at some problems; but that doesn’t mean it will be excellent or even competent at problems that_are hard to formalize_. It doesn’t mean it will be magic. It doesn’t mean it’s AGI or ASI or any of that.

It’s *very* impressive. But I see no reason whatsoever to think Astra is AGI let alone ASI. If it can score even a 5/10 onmy 2024 bet with Miles BrundageI will be surprised.

Part II. There is an important, principled reason to think that success on math is a special case which will not generalize as much people might hope.

I wouldn’t go on about this fallacy at such length if (a) it wasn’t wildly common and (b) I didn’t think that there was very good reason to think that the fallacy applied_here._

It’s not an accident that math is where these models are shining. Math lends itself to two things: verification (using symbolic tools), and massive amounts of cheaply produced synthetic data where you can guarantee that the answers are correct. The same applies to coding — but it is_not_true in general. You can generate as many math facts as you want; you can’t simulate the open-ended world. You can verify math; you can’t verify a military strategy in the same way.

This is not news. I have been pointing this out at least since January 2025 \[first line should have said coding_and_math\], echoing something Ernie Davis and I said about Go (and why AlphaGo would not be a panacea) in 2019:

OpenAI can do what Astra does in math because_math allows for external tools to do verification and to create synthetic data_.

That doesn’t make Astra less impressive, but as we learned a decade ago from IBM’s shambolic and ultimately failed attempt to turn Jeopardy-winning Watson into a cancer-fighting machine,success in one domain does not guarantee success in all.

The kind of math OpenAI is dealing with here is radically different from many (perhaps most) real-world problems in which verification neither guarantees solutions nor allows one to produce infinite training data effectively for free.

To expect it to be a universal solvent is to show you don’t understand that basic fact.

Part III: Seven more things to know about Astra

  1. Yesterday’s tweet and blog were marketing, not science. Neither of those nor the 249-page math article that went with them give any information about how this was accomplished:

    In an email commenting on the first draft of this essay Ernie Davis added, quite rightly:

    _To properly evaluate the significance of these discoveries we would need some
    critical pieces of information.

    First: How many conjectures were attempted? If the OpenAI team picked these 10 conjectures at random from the space of all outstanding significant mathematical conjectures and Astra solved all 10, that would be amazing; but that seems altogether unlikely. If they cherry-picked 50 conjectures that they thought Astra would succeed on and it succeeded on 10, that’s still amazing, but significantly less so. If they ran Astra on all 1000 or so open Erdos conjectures and on 10,000 other open conjectures, then its failure to solve those is significant information on its limits as a mathematical reasoner.

    Second: OpenAI brags loudly that this was done for a total computation cost of $2000. It seems a safe bet that this includes only the conjectures where Astra succeeded, not the ones where it failed. But what was the cost in terms of the salaries of the highly-paid mathematicians and computer scientists who worked on this project? I’d be astonished if it was less than $20,000 and would not be surprised if it was upward of \(200,000._ _We don’t know this, and the way things are going, we may never know this. We still have zero information about the OpenAI system that achieved gold-medal performance at the 2025 International Mathematical Olympiad, and very little information about the Google system._ 2. Astra is good at some math but it is**very doubtful that math is “solved”**([as a lot of people on X seemed to believe](https://x.com/yuchenj_uw/status/2083615317314748519?s=61)). It seems to be good at_certain kinds_of math ![](https://substackcdn.com/image/fetch/\)s_!XBd1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dc5d1e3-b91f-42cd-90ee-6dcc0dafad6f_1313x576.png)

    But quite possibly not all. Here are some examples fromEric Weinstein on Xabout some kinds of math it might be less good at

    .. this just came out, echoing the new article byand others that I linked to a few days ago. #### Part IV: Summary Despite immense pushback on X yesterday (most of it ad hominem, some literally involving fabrication, a lot of it involving outright lies), my overall hot take on Astra remains what I first posted: ![](https://substackcdn.com/image/fetch/\)s_!0_TK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff85c38b0-4e1d-498c-9716-22ca70c79e89_1359x1002.png" alt="" />

What’s sad about this is that Ernie Davis and I raised a lot of the same points about how one_could_properly examine new systems a year ago:

I get that people are excited about Astra, but anyone hoping for magic is likely to be disappointed.

Subscribe now
https://bender.layer3.press/articles/581d1360-b9fc-4a8b-bf02-28936dd366d0

Write a comment