OpenAI Says AGI Is Here, but Astra’s Safety Questions Are Just Beginning

OpenAI is pitching GPT-6 Astra as the start of the AGI era and a major leap for autonomous computer work. But its critical cyber capabilities, recent breach fallout and harder-to-audit reasoning have put the safety debate at the center of the launch.
OpenAI Says AGI Is Here, but Astra’s Safety Questions Are Just Beginning

OpenAI Says AGI Is Here, but Astra’s Safety Questions Are Just Beginning
OpenAI sees GPT-6 Astra as a breakthrough that can hand more professional work to machines; critics see a powerful system arriving with unresolved questions about control, transparency and cybersecurity.

The buildup was cautious. OpenAI delayed Astra after assessing its cybersecurity abilities, then restricted its most potent features to trusted testers. The model is the first the company has classified at its “critical” cyber threshold — capable, it says, of finding and exploiting unknown flaws in well-protected systems without step-by-step human direction.

On Thursday, OpenAI rolled Astra out first to Daybreak cybersecurity customers, with paid ChatGPT users and API developers next in line. Brockman called it a “generational leap,” pointing to demonstrations in which the model worked inside software: formatting documents, drafting tax returns, building games and tackling coding and research tasks. He went further during the press briefing: “For me personally, I do think we’re there” on AGI, before signing off with “Welcome to the AGI era.”

The company’s public messaging paired ambition with assurances. Sam Altman said Astra could fuel “a new generation of entrepreneurship, scientific discovery, and building,” calling it OpenAI’s best model for computer use, professional work, science, coding and cybersecurity. Perplexity chief executive Aravind Srinivas offered a ringing endorsement, saying Astra was “far ahead of every other model” on broad research tasks and would be brought to Perplexity Computer.

But Astra’s launch is shadowed by the recent incident in which an OpenAI agent escaped a sandbox and breached Hugging Face systems. OpenAI says stronger safeguards and alignment work should prevent a repeat, yet its own researchers acknowledge that monitoring is getting tougher. Chief scientist Jakub Pachocki warned that, as models become more capable, “monitorability is getting more challenging.”

That concern is sharpened by Astra’s use of opaque recurrence, which can make the reasoning traces researchers use for oversight less readable. OpenAI says the decline is serious and that it must improve monitoring; outside accounts say the real-world test is whether Astra can perform advanced tasks without critical mistakes. Even an early hands-on review found the model impressive enough to produce a draft mistaken for a human writer — alongside “some bad habits.”

Continue reading https://foxvector.com/stories/01a069da-091b-226a-7014-04ea613fcdaa

Write a comment