BusinessIssue #115

Half Your Salary on AI Tokens? Uber Found Out

What happened after Uber burned through a full year's AI budget in four months.

Half Your Salary on AI Tokens? Uber Found Out

Opening

This March, on the GTC stage, Nvidia’s Jensen Huang said something that stuck with me. “If an engineer earning a $500,000 salary only used $5,000 worth of tokens by year’s end, I’d go insane.” He said they should be spending at least $250,000 — half their annual salary. He went further, comparing skipping tokens to “designing semiconductors with paper and pencil.”

3 months later, this week, something unexpected happened at the company that followed that advice most faithfully. Uber capped AI coding tool usage at $1,500 per employee per month. The decision came after the company burned through its entire annual AI budget in just 4 months. To cut to the chase: “how much you spend on tokens” is already yesterday’s question. The next game is “what you spend it on” and “how you prove the results.”

🎯 Jensen Huang’s Token Doctrine

At this year’s GTC, Jensen Huang laid out a simple formula: “The more tokens you use, the more productive a developer becomes.”

In concrete terms: an engineer earning a $500,000 salary (about ₩700 million) should have an annual token1 budget of at least $250,000 (about ₩350 million). That works out to roughly $20,000 a month, or about ₩28 million. He went even further, proposing the token budget as a kind of fringe benefit stacked on top of salary — an additional payout equal to half of base pay, meant to make engineers 10 times more productive.

It’s worth noting the context behind this claim. Nvidia’s entire business model is, in effect, “a factory that produces tokens.” Huang himself laid out the formula at GTC: “Revenue = Tokens per Watt × Available Gigawatts.” The more tokens get consumed, the more GPU demand rises, and the more Nvidia’s revenue grows. For Jensen Huang, tokens are the product.

But for companies like Uber or Walmart, tokens are closer to raw materials. Not a product — a cost. Larry Dignan of Constellation Research put it precisely.

“JPMorgan sells financial services, Walmart sells retail, GM sells cars. What CIOs at these companies want is cheaper inference, better ROI, and a clear answer on when their AI investment pays off.”

The optimal strategy for those who sell tokens and those who buy them are exact opposites. And yet Huang’s declaration created a single culture across Silicon Valley: so-called “tokenmaxxing”2 — the belief that “consuming as many tokens as possible is productivity itself.” Some companies built leaderboards ranking individual developers’ token usage to spur internal competition. There were even reports that Meta ranked employees’ token consumption on an internal dashboard.

🔥 Uber’s Experiment, and the 4-Month Reckoning

Uber was tokenmaxxing’s most faithful practitioner.

In December 2025, Uber rolled out Anthropic’s Claude Code3 to roughly 5,000 engineers, alongside Cursor4. Straight out of Jensen Huang’s playbook, it built an internal leaderboard ranking employees by token usage. By the numbers alone, it was a resounding success.

By this February, usage had grown 2x, and 95% of engineers were using AI tools every month. 70% of code commits had shifted to AI-based generation, and the AI-powered feature adoption rate jumped from 32% in February to 84% in March. Last month, CEO Dara Khosrowshahi announced that “about 10% of all code is now written and submitted autonomously by AI agents.” He said AI usage was also rising rapidly in the legal and marketing teams. Uber had become what people call an “AI-native” company.

The problem broke in April. The entire annual AI budget had been consumed in just 4 months — a fact CTO Pravin Nepali Naga confirmed himself. Per-engineer monthly API costs ranged from $500 to $2,000, and multiplied across 5,000 engineers, that meant millions of dollars evaporating every month. Uber’s R&D spending was $3.4 billion in 2025 (up 9% year over year), but hit $950 million in Q1 2026 alone (up 17% year over year).

The CTO said the company needed to go “back to the drawing board.” And according to Bloomberg’s reporting this week, Uber introduced a $1,500-per-employee, per-tool monthly cap specifically for agentic coding5 tools. Because the cap applies independently per tool, employees could still spend $1,500 on Claude Code and another $1,500 on Cursor — but even combined, that’s just 7.5% of the roughly $20,000-a-month vision Jensen Huang laid out.

But the real bombshell wasn’t the cost. It was what COO Andrew Macdonald said last month on the Rapid Response podcast.

“I know our AI usage metrics are improving in astronomical ways. But it’s very hard to draw a straight line between any single one of those numbers and ‘we’re actually building features that are 25% more useful to consumers.’”

With 95% of engineers using AI and 70% of code written by it, this was an admission that the company couldn’t prove what value any of it delivered to customers. Uber’s stock fell 2.9% after the remark.

📐 The Missing Piece: Measurement, the Next Game

This isn’t just Uber’s problem.

This week, Walmart switched its internal AI agent, ‘Code Puppy,’ to a per-employee token rationing system. Amazon had run an internal AI usage leaderboard, but tore it down after employees started competing purely on token consumption instead of actual work output. One unnamed company reportedly set no spending cap at all and ended up racking up $500 million (about ₩700 billion) in AI token costs in a single month.

The scale explains why this keeps happening. According to Menlo Ventures’ analysis, 55% of corporate departmental AI spending is now concentrated on coding tools — up from $550 million in 2024 to $4 billion in 2025, a 7x jump in a single year. Gartner projects AI agent software spending will reach $207 billion in 2026 — up 139% year over year.

Money is pouring in. The infrastructure to measure its effect is not.

CloudBees’ “Code Abundance” report, released this year, makes the gap plain. 83% of corporate leaders self-rated their organizations as “well-prepared for AI adoption.” Yet at the same time, 81% said “production issues caused by AI-generated code have increased.” The gulf between confidence and reality is that wide.

Developer productivity benchmark research is reaching the same conclusion. The metrics companies used to rely on — weekly PR6 counts, lines of code, commit counts — are no longer trustworthy in 2026, because AI inflates code volume. You can tell whether quantity increased, but the old dashboards can’t tell you whether value increased. AI coding tools generate genuine productivity gains, but they also generate the illusion of productivity gains — and the core problem is that existing metrics can’t tell the two apart.

Oz’s Lens

Honestly, I expected something like this the moment I heard Jensen Huang’s token doctrine.

There’s a pattern I’ve watched play out repeatedly in developing technology strategy. Whenever a new tool arrives, companies always follow the same sequence: enthusiastic adoption → cost shock → only then does ROI measurement begin. It happened with CRM. It happened with cloud. Now AI coding tools are walking the same path. Measurement always comes last.

This time, there’s a structural reason measurement is especially hard. Writing code is just one part of the software development pipeline. Even if AI makes the coding segment 10 times faster, if the bottleneck sits in design or testing, overall speed doesn’t change. It’s the same principle as manufacturing’s “Theory of Constraints.”7 And on top of that, using a lot of tokens means using AI “a lot” — not necessarily using it “well.” The number of calls you make and the number of deals you close are different metrics.

Still, it’s hard to deny the value of AI coding tools altogether. Uber didn’t stop using them — it capped them. The end of tokenmaxxing doesn’t mean the end of the AI tools era. It means the era of measurement has just begun.

Closing

Jensen Huang said “use more” (thesis), and Uber followed faithfully, burning through its annual budget in 4 months (antithesis). What companies are realizing now is that the real competitive edge isn’t how much you use tokens, but your ability to prove the link between AI and business results (synthesis). Measurement design has to come before adoption speed.

How does your organization measure the impact of AI tools? And have you experienced your own version of “tokenmaxxing”? Let’s talk about it in the comments.

References & Further Reading

Primary sources

Background

The author, Kwangseob Ahn, is a professor of business administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — business data management and business analytics — while leading GTM and AI strategy consulting in the field, designing the seam between technology and business. He has published academic research on a memory architecture for AI dialogue systems (HEMA) and runs Daily Arxiv, a daily curation of global AI papers. He holds a master’s from Korea University’s Graduate School of Technology Management and a KMBA. He is the author of Homo Brainless: The People Who Outsource Their Thinking.

Footnotes

  1. Token: The basic unit AI uses to process text. Roughly equivalent to one Korean character or 0.75 English words. AI service costs are typically billed based on token consumption.

  2. Tokenmaxxing: A coined term for the corporate culture of maximizing AI token usage. Under the assumption that “more usage means more productivity,” some companies even ran usage leaderboards.

  3. Claude Code: An agentic coding tool built by Anthropic. It goes beyond simple code suggestions, autonomously writing, editing, and executing code directly in the terminal.

  4. Cursor: A development tool that integrates AI directly into a code editor. Built on VS Code, it lets AI analyze and edit code across entire files.

  5. Agentic Coding: A coding approach where AI goes beyond suggesting a single line of code to autonomously planning and executing multi-step tasks — creating files, running tests, and fixing errors on its own.

  6. PR (Pull Request): The process by which a developer submits code for team review before it’s merged into the main project.

  7. Theory of Constraints: The theory that a system’s overall performance is determined by its slowest bottleneck. Speeding up just one stage doesn’t change the overall pace if another stage remains blocked.