Pause OpenAI, now

Quite simply, they can no longer be trusted.
Pause OpenAI, now

I have often counseled calm where others might counsel panic.

I told you that the Hugging Face incidentcould likely have been prevented had best cybersecurity practices been followed. (And I stand by that.)

I told you (and most people still seem unaware) that that the OpenAI Hugging Face incident was part of a training exercise, with some internal guardrails shut down, so it was not quite as bad as it seemed1

I told you that Astra probably wasn’t AGI.2

And I stand by all of that.

But I am freaked out.

What I am freaked about is not imminent AGI.

It’s OpenAI.

I simply don’t believe that they are trustworthy enough or responsible enough to be good stewards of the technology that they are developing.3As a company, they simply don’t have good judgment.

§

Here are four considerations.

  1. Sam Altman cannot be trusted.I have been writing about that for a long time. Ronan Farrow’s reportingbacks that up. So does the just-dropped bombshell below that I am about to get to.

  2. The just-released Astra reduces Chain fof Thought (CoT) monitorability, one of the few (not especially reliable, but better than nothing) tools we have for keeping generative AI from running wild. The AI safety community is up in arms about this—with good reason. There are tons of posts like this now, all quite right:

    The decision to release Astra is a clear example of the willingness of OpenAI management to trade off safety in exchange for relatively modest gains in performance. Thered alert that I sounded a couple days agowas on target. They really_are_playing around with new techniques that reduce monitorability. Andtheir own data shows that monitorability is in fact compromised to some degreein the newly released Astra, particularly on “destructive actions.” They released it anyway. That speaks volumes.

  3. Something I read last night, and that only fully clicked into place this morning (see fact 4 below) terrifies me. A prominent recently departed employee (who presumably still owns significant stock, and who has repeatedly struck me as an advocate of OpenAI since he left) wrotean essay on Xbasically asking people to simply accept that rogue AI is here to stay.

§

How Trump handles OpenAI may end up defining his legacy. I estimate the probability of a major cyber incident attributable to OpenAI in the next 12 months to be very high, certainly over 50%. And Trump, if he doesn’t intervene, may share some of the blame.

In the meantime, I call upon Congress to investigate OpenAI, with an eye to whether the company might need to be sanctioned—or even paused— now.

Please share this call to Pause OpenAI

Share

Subscribe now

1

FromOpenAI’s own report on the Hugging Face incident“The incident occurred during routine testing”, and “At the time of the incident, OpenAI estimated maximal cyber capabilities by running this evaluation without the production classifiers intended to prevent models from pursuing high-risk cyber activity”. Had those classifiers been turned on, the incident might well not have happened.

2

Data from Epoch AI tends to support this conclusion. Astra is genuine improvement but not even statistically off of trend; if it really were AGI I think we would expect to reflect a sharper departure from previous models.

3

Altman told Alex Heath the other day that “Altman wants OpenAI to be seen as “the most responsible company … good stewards of technology”. As made plain in this essay, he’s talking the talk, but not walking the talk. We do desperately need good stewards. He’s right about that. Unfortunately Altman himself is manifestly not suited to that particular job.
https://bender.layer3.press/articles/caaa8b39-455a-4adf-9618-faee0caf5db9

Write a comment