Google’s New Flash Model Looks Cheap—Until the Token Bill Arrives

Google is pitching Gemini 3.8 Flash as a fast, powerful and low-cost leap for coding and cyber defense. But the unchanged per-token price may conceal higher real-world spending as the model takes more reasoning steps.
Google’s New Flash Model Looks Cheap—Until the Token Bill Arrives

Google’s New Flash Model Looks Cheap—Until the Token Bill Arrives
Google sees Gemini 3.8 Flash as proof that rapid iteration can deliver frontier-grade AI at Flash-level prices. Early outside analysis, however, suggests that a model built to think and act more may turn its headline bargain into a larger bill.

Google released Gemini 3.8 Flash just weeks after version 3.7, its third Flash update in six weeks. The company says the model “works harder,” taking more reasoning steps on difficult tasks and repeatedly calling tools, while retaining introductory API rates of $0.75 per million input tokens and $3.75 per million output tokens.

That sticker price is central to Google’s pitch. CEO Sundar Pichai said 3.8 Flash delivers “significant leaps” in software engineering, agentic tasks and multi-step reasoning, with DeepSWE v1.1 results that he said surpass most larger frontier models. Another account of the release described it as Google’s strongest reasoning and coding Flash model yet, while noting that its gains over 3.7 are more pronounced in coding than in many other tests.

But the economics are less straightforward. Google cautions that 3.8 Flash may consume more tokens to maximize performance, particularly at higher effort settings. Artificial Analysis estimated the effective cost is about 40% higher than Gemini 3.7 Flash, despite identical token rates, because output per task rose 30% and agentic evaluations required more turns. Developers focused on minimizing consumption can continue using the earlier model.

Alongside the standard release, Google introduced Gemini 3.8 Flash Cyber, tuned to find vulnerabilities and produce patches. Pichai called it the company’s “most capable cybersecurity model,” citing frontier-level vulnerability discovery and patching at Flash speed and pricing. Access is narrower, however: the cyber model is limited to governments and trusted partners through Google’s Fairwind Program, even as the standard Flash model rolls out to subscribers, developers and enterprise users.

The message is clear: Google is selling more capability for the same listed rate. The catch is that more capable AI may also be more expensive to run.

Continue reading https://foxvector.com/stories/01a0640e-76d1-01c6-716c-0b8c47432a68

Write a comment