Red Alert: OpenAI is poised to cross an AI safety redline.

Making models harder to monitor is not what we need
Red Alert: OpenAI is poised to cross an AI safety redline.

The Information just broke the scoop thatOpenAI is playing around with a new technique, in which models will reveal less of their “thinking”, making them harder to monitor.

The new techniques they are exploring may make such monitoring difficult or impossible.

§

Last year, an all-star cast wrote a fascinating paper that feels deeply relevant now, called_Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety_, I fully agree with the highlighted bit:

They are exactly right. CoT monitoring is imperfect (asSubbarao Kambhampatiand others have shown), but it is one of the best threads we have for monitoring the giant black boxes that we call LLM. It is a slender thread, but sacrificing it thread for (small?) performance gain feels like a dangerous game.

Earlier tonight Steven Adler, of Guidelight.ai andone of the many researchers to have departed from OpenAI’s safety teams, said this, echoing Nathan Calvin:

I fully, 100% agree.

Share

Subscribe now

P.S. Bonus those who prefer Terminator references:
https://bender.layer3.press/articles/112c4407-ff8f-4e63-b198-c761ba4e35b2

Write a comment