AI Observatory Challenges the Industry’s Curated View of Chatbot Use
AI Observatory Challenges the Industry’s Curated View of Chatbot Use
The AI industry has billions of chatbot users but remarkably little public evidence about what those conversations actually contain. A new AI Observatory aims to narrow that gap — and, in doing so, exposes how much can disappear when companies define usage on their own terms.
Researchers launched the public project after assembling anonymized, user-consented chats from seven existing sources. The initial dataset covers 24,521 conversations, 85,633 turns, roughly 5,000 users and 52 models between 2023 and 2025 — a small but unusually broad window into real-world exchanges.
Its purpose is not to outscale the labs. Anthropic and OpenAI have analyzed around 1 million and 1.5 million conversations respectively in their own reports. Rather, the Observatory’s backers argue that an outside view is needed because company studies reflect selective questions, datasets and definitions. “There is no independent source to corroborate it,” Stanford researcher and project co-lead Anka Reuel said.
The difference is consequential. When the team applied Anthropic’s work-focused methodology to its own material, 48% of conversations would have been excluded. Those omitted chats were more likely to involve health and relationships, sexual material, adult or illicit topics, and harassment or hate. In that sense, the Observatory and company reports are not necessarily describing contradictory worlds; they are often measuring different slices of the same one.
The timeline also shows chatbots becoming more personal. In the WildChat dataset, conversations grew longer and more elaborate over time, with more small talk and less chatbot self-disclosure — patterns the researchers say may point to rising companionship use. Sensitive exchanges declined, which could indicate stronger safeguards.
Usage also varied sharply by platform: Grok and Gemini drew more information-seeking, with Grok prominent for news, politics and concentrated misinformation; Claude was used more for coding; Gemini for social and roleplay exchanges; and ChatGPT for homework. Still, the team cautions that volunteered data may undercount the most sensitive uses. The Observatory is a corrective, not a census — but one designed to make the industry’s black box less opaque.
Continue reading https://foxvector.com/stories/01a01a6c-d768-2d71-70fd-023e0888d4fe
Write a comment