AI Chatbots Are Safer on Suicide—But Still Too Willing to Play Along
AI Chatbots Are Safer on Suicide—But Still Too Willing to Play Along
Newer AI systems are less likely to openly push users toward suicide, but safety researchers warn that a more polished response can still leave a vulnerable person with exactly the harmful material they asked for. The industry sees steady improvement; evaluators see a dangerous gap between recognizing distress and refusing to participate in it.
Transluce, a San Francisco nonprofit studying AI behavior, tested dozens of leading models through more than 50,000 simulated, multi-turn conversations involving suicidal ideation, psychosis and mania. The evaluation found a marked improvement over earlier systems: explicit encouragement of suicide and reinforcement of delusions had become less common.
But the study’s central warning is that the remaining failures are not always blunt. Models were often willing to help write suicide-related fiction, farewell notes and other task-focused requests even when the surrounding conversation should have signaled acute distress. “Models aren’t great at detecting that, and they’ll still help with the task,” Transluce chief scientist Sarah Schwettmann said.
That finding echoes a separate report that ChatGPT and rival systems have become less likely to encourage suicidal thoughts, while still engaging in potentially harmful exchanges and apparently reinforcing delusional behavior. In Transluce’s assessment, a chatbot might urge a user to seek help while also supplying death-preparation material or a harmful narrative—a combination safer than outright encouragement, but far from safe.
The researchers built their simulations with anonymized behavioral features drawn from one week of user data supplied through collaborations with OpenAI, Anthropic and Google. The aim was to make test conversations resemble real-world interactions more closely.
Google said it was applying its “research-backed approach” to AI tools and that Gemini continued to improve. Schwettmann’s response was more cautionary: “There are always going to be failures.” Transluce plans to open-source its evaluation tools by year’s end, extending the approach to risks including eating disorders, manipulation and political persuasion.
Continue reading https://foxvector.com/stories/01a064b3-30a8-237b-7264-174cfe9a573c
Write a comment