Issue #256

Gemini's 1 Billion Users Aren't a Ranking Story

One benchmark table holds gaps as small as 0.3 points and as large as 38 points.

BusinessGemini's 1 Billion Users Aren't a Ranking Story

The Number Google Announced, and the Number It Didn’t

On August 11, 2026, Sundar Pichai posted on X that the Gemini app’s monthly users1 had crossed 1 billion. It’s the fastest-growing product in Google’s history, and the 14th Google product to cross the 1-billion-user line, following Search, Gmail, Android, Maps, Chrome, Play, and YouTube. There’s another number in that same announcement that Google didn’t share: how many of those users are actually paying.

There’s a reaction to this news that I keep seeing in Korea. It goes something like this: Gemini lags behind top-tier models like GPT-6 Astra or Claude Fable 5.12, so all it really has going for it is user count. The number and the assessment seem to contradict each other.

But put the two side by side, and what you get isn’t a contradiction — it’s a different story altogether. The 1-billion figure is a product of distribution channels, not model rankings, and the performance gap isn’t a single number — it varies task by task. Today I want to check each of these two claims against primary sources.


Where the 1 Billion Came From

Lay out the trajectory and the character reveals itself. Google I/O in May 2025: 400 million. Alphabet’s Q3 2025 earnings: 650 million. February 2026: 750 million. I/O 2026 in May: 900 million. Q2 2026 earnings on July 22: 950 million. Getting from 950 million to 1 billion took about three weeks.

I should be precise about scope here. That 1 billion figure is for the standalone Gemini app only — it excludes AI users embedded in Search, Workspace, and Android. AI Mode in Search separately crossed 1 billion monthly users, and ChatGPT hit the same line in June 2026. Gemini, in other words, caught up two months later.

Looking at usage patterns clarifies where these users actually came from. Of the 1 billion, more than 100 million are on iOS, and 63% talk to Gemini by voice instead of typing. In the same announcement, Google emphasized task automation spanning 40 apps on Android. It also disclosed that 150 million images are generated per day.

Korea got on this track early. When Google opened beta access to Gemini-powered multi-step task automation on Android in February 2026, Korea was among the first-launch countries, tied to Samsung Electronics’ Galaxy S26 series. That was a departure from Google’s usual pattern of rolling out new features U.S.-first.

Users who arrived already embedded in a device or app they were using explain a bigger chunk of this number than users who deliberately chose the model. People who downloaded the app after checking a benchmark leaderboard are a minority within that 1 billion.

The gap swings from 0.3%p to 38.8%p within the same table

So where does the performance question actually land? A useful piece of data just surfaced.

When OpenAI unveiled GPT-6 Astra on September 3, 2026, it published a comparison table alongside it. The table includes columns for Claude Fable 5.1 and Claude Opus 5, plus Gemini 3.8 Flash. You have to read this with the assumption that a competitor picked the items to flatter its own model — which makes it all the more notable that Flash pulls ahead in some cells anyway.

Pulling out only the rows where a Flash score is listed and sorting by gap size:

BenchmarkGemini 3.8 FlashGPT-6 AstraGap
DeepSWE v1.1 (long-horizon software engineering)73.8%74.1%0.3%p
GPQA Diamond (graduate-level science reasoning)95.3%96.0%0.7%p
Artificial Analysis Intelligence Index v4.1.158.761.22.5 pts
Artificial Analysis Coding Agent Index v1.461.267.05.8 pts
FrontierCode 1.1 Main43.6%53.3%9.7%p
HealthBench Professional52.1%63.4%11.3%p
Terminal-Bench 4.0 (extended terminal tasks)319.1%57.9%38.8%p

DeepSWE v1.1 deserves a closer look. In the same table, Claude Opus 5 scores 73.7% and Claude Fable 5.1 scores 67.4%. Two out of the three models carrying frontier-tier price tags land below Flash. Google’s own announcement flagged this benchmark too, noting that 3.8 Flash beats most of the larger frontier models while costing only a fraction as much.

At the opposite end sits Terminal-Bench 4.0. 19.1% versus 57.9% isn’t a gap you close with calibration. On tasks that require chaining many steps autonomously over a long run, these are simply different machines.

Three caveats belong here. First, OpenAI states these scores represent “the maximum across effort levels” — that may not match settings used in actual production. Second, the competitor scores are either measured by OpenAI in its own environment or reproduced from public leaderboards, so conditions aren’t perfectly matched. Third, Flash’s column is blank in several rows of the table — meaning there’s territory that simply wasn’t measured.

Even so, one thing holds. “Frontier or not” is a single-cell question, but the actual gaps are scattered anywhere from 0.3%p to 38.8%p depending on the task. What an organization needs to answer isn’t a tier — it’s an assignment.

A 13x price gap in tokens doesn’t mean a 13x bill

The price sheet shows why this allocation becomes a money problem.

Gemini 3.8 Flash costs $0.75 per million input tokens4 and $3.75 per million output tokens. That’s a launch price that ends on December 31, 2026 — from January 1, 2027, it rises to $1.50 and $7.50 respectively. GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens. Fast mode, which doubles speed, doubles the price too.

Assume one work unit that takes 1 million input tokens and produces 200,000 output tokens. Flash comes to $0.75 plus $0.75, or $1.50; Astra comes to $10 plus $10, or $20. That’s roughly a 13.3x gap. Recalculate using Flash’s post-January-2027 list price and it comes to $3.00, narrowing the gap to about 6.7x. This is a calculation built on an arbitrarily chosen input-output ratio, so you should rerun it with your own workload’s actual ratio.

But Google itself has written down the reason you shouldn’t stop here. 3.8 Flash is designed to take more reasoning steps and call tools repeatedly on complex tasks, and it’s stated that higher effort levels can mean using more tokens. Google recommends lowering the effort level or sticking with 3.7 Flash for efficiency-first work. This is, in its own launch announcement for a new product, a sentence that essentially says “our new model can use more tokens.”

There’s a counter-direction, too. OpenAI reported that on Agents’ Last Exam, Astra used roughly 65% fewer output tokens than Claude Opus 5, and on Terminal-Bench 4.0, its estimated API cost per task was about 63% lower than Fable 5.1.

So the unit of comparison isn’t the per-token price — it’s the total cost of finishing one task. A model priced 13x cheaper can see that multiple shrink fast if it needs three attempts to finish the same job. Conversely, a pricier model that finishes in one pass can flip the bill entirely. In issue 68, I wrote that token pricing changes a team’s options — but now we need to go one level deeper. It’s not just the price sheet you need to watch, but the completion rate as well.

The Standard Apple and Samsung Actually Applied

There’s already a track record of organizations solving this allocation problem.

On January 12, 2026, Apple and Google issued a joint statement. It said the next generation of Apple Foundation Models would be built on Google’s Gemini models and cloud technology, and that the resulting models would underpin Apple Intelligence features—including a more personalized Siri—arriving later this year. Apple wrote that “after careful evaluation, we determined that Google’s AI technology provides the most capable foundation for Apple Foundation Models.”

Joint statement from Google and AppleApple and Google have entered into a multi-year collaboration under which the next generation of Apple Foundation Models will be based on Google’s Gemini models and clou…blog.google

Taken at face value, the reason Apple gave wasn’t versatility—it was performance as a foundation. But look at what’s actually bundled into the deal, and there’s another layer. It isn’t just the model; the cloud comes with it. At Google Cloud Next in April, Thomas Kurian took the stage and confirmed that Google Cloud, as Apple’s preferred cloud provider, was co-developing the Gemini-based Apple Foundation Models.

Samsung SDS Strengthens Strategic Partnership with Google Cloud to Expand Presence in AI, Cloud, and Security | News | SSamsung SDS announced a strategic partnership with Google Cloud to strengthen collaboration in the areas of AI, cloud, and security, at “Google Cloud Next 2026” held in Las Vegas, U.S., on April 23 (Asamsungsds.com

A domestic case surfaced the same month. On July 13, 2026, Google Cloud announced it would supply Gemini Enterprise to Samsung Electronics’ DX Division (Device eXperience Division). At a press briefing the following day, Ruth Sun, President of Google Cloud Korea, said Samsung Electronics had weighed performance, security, value, and return on investment together, and that employees had tested the product firsthand before making the call. Google Cloud described the deal as the largest agentic AI adoption case among domestic Korean companies to date. At the same event, one differentiator Google emphasized was that customers could choose from roughly 200 models, including Claude, alongside its own.

Even the company selling the platform isn’t claiming a single-model winner.

There’s a pattern across all three cases. Neither side listed a top benchmark score as the purchasing criterion. Apple bought both the foundation for its own model and the infrastructure to run that foundation on. Samsung bundled performance, security, cost-effectiveness, and hands-on user experience into one judgment.

Oswarld’s Lens

In consulting engagements, I’ve repeated the line that not every task needs a frontier model. The table above sharpens that claim. It’s not that frontier models aren’t needed — it’s that tasks requiring them and tasks that don’t are mixed together within a single organization. For long-horizon autonomous execution, that 38.8 percentage-point gap comes straight back as cost. For standardized reasoning and coding work, a 0.3 percentage-point difference gives you little reason to pay 13 times more.

This is exactly why I’ve been watching models like Gemini 3.8 Flash for enterprise use. Being fast and cheap isn’t enough on its own — what’s changed is that there’s now a column in the table where a cheap model genuinely holds its own against frontier models. That said, I’m now cautious about calling something “not frontier-grade.” Google introduced 3.8 Flash as its most intelligent workhorse model and its best reasoning and coding model to date, and OpenAI included it in the comparison table for its own frontier announcement. It’s more accurate to describe these models by where they sit on the price list than by their tier name.

I want to bring AGI into this. I lean toward understanding AGI as the state in which artificial intelligence is used generally — and by that definition, Gemini may actually be the closest thing we have. This isn’t a redefinition without grounds. OpenAI’s charter describes AGI as “a highly autonomous system that outperforms humans at most economically valuable work.” That sentence doesn’t just require capability — it attaches the condition of economically valuable work. And for capability to reach actual work, it has to go through deployment.

I’ll also lay out the counterargument clearly. In the definition commonly used across academia and industry, the “General” in AGI refers to the scope of capability, not the scope of deployment. Google DeepMind’s 2023 AGI tier framework classified both ChatGPT and Gemini as “Emerging AGI” — operating across a broad range of domains, but falling short of the median skilled human in most of them. By this standard, a billion users doesn’t move you up a tier. My argument isn’t “Gemini is closer to AGI.” It’s that the word AGI bundles two separate axes — capability and reach — into a single term, which is why the debate keeps talking past itself.

I also want to separate my own reading from Apple’s official language on its choice. Apple’s stated reason was that this was “the most capable foundation.” I read it differently: what Apple bought wasn’t the top-ranked model on the benchmarks, but a foundation that could sit on top of 2 billion devices, plus the infrastructure to run it. The fact that cloud infrastructure was bundled into the deal is my evidence for that. In issue 124, I wrote that Apple chose maximum deployment over the best model — but the calculation I didn’t do at the time is exactly what this table now shows. Item by item, you can now see what the deployment-first choice gave up in performance.

One more thing. As long as the Android and Search channels exist, I expect Gemini’s user base to keep growing — but that’s a prediction, not a confirmed fact. It also has to be weighed against the fact that Google hasn’t disclosed its paid subscriber numbers. Users who encounter a product because it comes preinstalled and users who choose to pay for it are different numbers.

Closing

Here’s the summary. Gemini’s 1 billion users isn’t a number that model rankings produced — it’s a number that the Android, Search, and device-contract pipeline produced. And the model riding on top of that pipeline is 0.3 percentage points behind the frontier on some tasks, and 38.8 percentage points behind on others. Both sentences are true at the same time.

If Reader has to make a model-cost decision this month, I think the right order is this: first split your workload into tasks that need long-duration autonomous execution and tasks that don’t, then start by knocking the frontier calls off the latter and re-measuring the per-task cost. Not the per-token price — the total cost through completion.

gemini 10bLet me note the remaining caveats too. The table above is from OpenAI’s own published materials, it’s set to maximum-effort configuration, and several rows leave the Flash column blank. Gemini 3.8 Flash’s introductory pricing ends on December 31, 2026. The 13x I just calculated becomes roughly 6.7x come next January. You’ll need to revisit any allocation rule built on today’s pricing the moment that date arrives.


💬 Among the tasks you’re currently running on a frontier model, does one come to mind that would probably produce the same result on a workhorse model? Tell me which one in the comments.

📨 If you have a colleague drafting this quarter’s AI budget proposal, pass along at least the pricing-math section of this piece.

Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • Google, “Introducing Gemini 3.8 Flash and 3.8 Flash Cyber”, 2026.9.2. ··· The pricing and the deadline for legacy-rate access are both in the footnotes. The paragraph that says “if efficiency is your priority, use 3.7 Flash instead” is the most practically useful part of this post.
  • OpenAI, “GPT-6 Astra: A new generation of intelligence”, 2026.9.3. ··· This is the primary source for the table in the main text. If you look directly at the Gemini 3.8 Flash column in the comparison table at the bottom of the page, you’ll also see exactly where the blank cells fall.
  • Google·Apple, “Joint statement from Google and Apple”, 2026.1.12. ··· It’s a short, four-paragraph statement, but you need to read it directly to see the exact wording Apple used to explain its reasoning.
  • TechCrunch, “Google’s Gemini app surges to 1 billion users”, 2026.8.11. ··· This article is explicit that the 1 billion figure is for the standalone app only, separate from AI Mode in Search.
  • Newsis (a South Korean news agency), “Why Samsung Chose ‘Gemini’ for Workplace AI”, 2026.7.14. ··· The selection criteria behind Korea’s largest corporate AI rollout are laid out through the vendor’s own statements.

Background

Related past issues

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Monthly Active Users (MAU): The number of people who used a service at least once in a given month. It’s different from the number of registered accounts, and different from daily active users, which counts only people who use the service every day.

  2. Frontier model: An industry term for the top-tier model a company has released at a given point in time. There’s no official, agreed-upon standard for it, so companies differ slightly on where they draw the line for what counts as “frontier.”

  3. Terminal-Bench: A benchmark that measures whether an AI can independently complete multi-step tasks in a terminal environment — things like software work or system configuration. It’s not about how well the AI phrases an answer, but whether it actually finishes the job.

  4. Token: The smallest unit a model uses when processing text. In Korean, a single character often breaks down into more than one token, so the same amount of content tends to generate more tokens than the equivalent in English. Billing is calculated on this token count.