Buying GPUs Isn't Enough: AI Infra's Real Battleground
The real bottleneck in AI infrastructure isn't chips — it's whoever locks up power and memory first who wins.

Opening
Dear reader, let me start with one number: ₩690 trillion (~$496B).
That’s the amount five Big Tech companies have pledged to pour into infrastructure in 2026 — roughly 30% of Korea’s annual GDP. And yet there’s a company that can’t spend all that money even if it wants to: Microsoft. CEO Satya Nadella admitted it himself — GPUs are piling up in warehouses with no power to plug them into. There’s $80 billion worth of backlogged Azure1 orders, and physically, there’s no way to run the servers.
In the last week of February 2026, Big Tech announcements came in waves: a Meta-Nvidia GPU partnership on February 17, a Meta-AMD deal on February 24, and Google’s Texas data center announcement the same day. Most coverage focused on “another GPU deal.” I think these announcements are telling a different story.
That’s what today’s issue is about — not who bought the most GPUs, but where the real bottleneck actually lies.
The Structure of the AI Infrastructure War: Three Axes
Building AI infrastructure requires three things at once.
The first is high-compute chips — AI-specific semiconductors like GPUs and TPUs2. The second is power infrastructure — the data centers and electricity needed to actually run those chips. The third is consumable memory chips: semiconductors like HBM3, DRAM, and NAND that hold data while AI models are computing.
The problem is that no company currently dominates all three at once. So what’s actually happening is this: everyone is pouring in ₩690 trillion (~$496B) while racing to shore up whichever axis they’re weakest on.
Let’s break this down one axis at a time.
Axis 1: High-Compute Chips — A Market With Only Three Options
Externally sourceable high-compute AI chips essentially come from just three places: Nvidia GPUs, AMD GPUs, and Google TPUs. It’s what’s commonly called an oligopoly4. And it’s a fairly structural one at that. From the software ecosystem (Nvidia’s CUDA5) to packaging technology to relationships with TSMC, entering this market as a new player is a multi-year undertaking.
That’s why Big Tech is now formalizing supply diversification strategies. Meta is the clearest example. <On February 17, it signed a long-term partnership with Nvidia for millions of Blackwell and Rubin GPUs, and a week later, on February 24, it signed a multi-year deal with AMD for up to 6GW of Instinct GPUs.> At the same time, it’s also developing its own MTIA chips.
Mark Zuckerberg called it “an important step to diversify our compute.” That’s an official declaration that you can’t rely on just one source. AMD CEO Lisa Su described the deal as “a multi-year, multi-generational collaboration to deliver high-performance, energy-efficient infrastructure optimized for Meta’s workloads.”
Google is moving in a different direction. It’s doubling down on an inference6-focused chip strategy with its own TPU v7 Ironwood, signing a deal to supply Anthropic with over a million TPUs and starting to sell to external neoclouds as well. In other words, Google isn’t just using TPUs internally anymore — it wants to sell them the way Nvidia sells GPUs.
But let’s be clear about one thing here: Nvidia isn’t collapsing. It’s still dominant in general-purpose flexibility and its software ecosystem. It’s just that for large-scale inference workloads7, custom chips like TPUs and ASICs8 have started to pull ahead on cost efficiency. The market is shifting from a monopoly to a layered structure.
Axis 2: Power Infrastructure — “We Have the Money, Just Not the Electricity”
In my view, this is the most physical and the most serious bottleneck of the three.
On February 24, Google announced a new data center in Wilbarger County, Texas, saying it had contracted over 7,800MW of net new power on the Texas grid. Since 1GW roughly powers 700,000 households, 7,800MW exceeds the electricity used by the combined populations of Busan, Incheon, and Daegu — three of Korea’s major cities. This also includes a joint clean-energy build-out with AES and air-cooling systems designed to minimize water use.
Meta is building a 1GW Prometheus data center in Ohio and planning up to 5GW for Hyperion in Louisiana. Zuckerberg himself mentioned that the site is about 20 times the size of Yeouido, Seoul’s financial district island.
Why do they need so much? Because AI is structurally driving up power demand. The U.S. Energy Information Administration (EIA) projects U.S. commercial power demand will grow 3% in 2025 and 4.5% in 2026, while the International Energy Agency (IEA) expects global data center power consumption to double by 2030.
But here’s the real problem: lead times for power transformers have stretched to 128 weeks (about 2.5 years). Even if you order one today, you won’t get it until the second half of 2028. Money can’t buy you out of a two-and-a-half-year wait.
That’s why Texas is emerging as the go-to data center location — it has its own independent grid, ERCOT9, flexible regulation, and vast land. Google’s choice of Wilbarger County isn’t just about land — the key is securing electricity in advance. We’re already seeing what happens to companies like Microsoft that haven’t secured enough power: their GPUs are sitting idle in warehouses.
There are even reports that the ambitious Stargate Project10 — backed by SoftBank, the U.S. government, and OpenAI — has scrapped its original plan to concentrate construction in Texas due to the power crunch, opting instead to build smaller or split across multiple sites.
Axis 3: Memory — AI Is Devouring Memory Supply
This third axis is the one most closely tied to Korea.
2026 HBM production is already essentially sold out. In its October 2025 earnings call, SK hynix said its 2026 HBM, DRAM, and NAND production capacity was “essentially sold out,” while Micron announced it would exit the consumer memory market entirely to focus solely on AI data center customers.
Why did this happen? There’s a structural reason.
HBM3E requires roughly three times the wafer area of standard DDR5. Attaching memory to a single AI accelerator consumes the equivalent of three ordinary memory production lines. The more memory makers focus on HBM production, the less ordinary memory is left over for the smartphones and laptops we use.
The results are already visible. Samsung raised the price of its DDR5 32GB modules from $149 to $239, an increase of about 60%. DDR5 contract prices have surged more than 100%. There’s even talk that the consumer memory shortage could delay Sony’s PS6 launch (this isn’t officially confirmed by Sony — it’s a rumor from console gaming trade press).
So how long will this last? SK hynix executives, speaking at a JPMorgan meeting, expressed confidence that the memory upcycle could continue into 2027, and possibly through 2028–2029. They expect shortages to persist even as CSPs11 place double and triple orders to hedge against scarcity.
There’s one more thing worth flagging: leverage in long-term agreements (LTAs) has flipped. In the past, when memory was in oversupply, suppliers were the ones pushing contracts. Now, cloud companies like Google and Microsoft are the ones coming to suppliers, asking them to lock in volume. That’s a signal that memory has been upgraded from a commodity component to a strategic asset.
The structure of memory demand is also shifting as AI enters its inference era. Demand used to be concentrated in training; now the center of gravity is moving toward inference. As KV cache12 emerges as a new driver of memory demand, the amount of DRAM needed per server is increasing, pushing up demand for enterprise SSDs (eSSDs) as well.
BofA projects the 2026 HBM market at $54.6 billion, up 58% year-over-year, and named SK hynix its top pick in global memory.

Oz’s Lens
Honestly, watching this competition reminds me of a pattern I saw often back when I worked as a GTM strategist.
It’s the pattern of a moving bottleneck. Whenever a product overcomes one constraint, the bottleneck always shifts somewhere else. As AI chip supply increased, power became the constraint; once power was secured, memory became the new one.
From this angle, which company is best positioned right now? I’d say Google. With its own TPU ecosystem, 7.8GW of power secured in Texas, and an AI-stack partnership with Anthropic, it has the most balanced portfolio across all three axes.
Meta chose a smart strategy of chip diversification — a three-way split across Nvidia, AMD, and MTIA. But that spreads supply risk; it doesn’t actually solve the bottleneck. The real variable for Meta is when its 5GW Hyperion facility in Louisiana actually comes online.
Microsoft is in the toughest spot right now. It’s heavily dependent on Nvidia and is furthest behind on securing power. The fact that $80 billion in unmet orders keeps piling up means it’s missing out on monetization opportunities at this very moment.
And we can’t skip the Korea story here. SK hynix and Samsung are the key players dominating the memory axis of these three. SK hynix holds roughly 50–62% of the HBM market, and that share is likely to hold up in HBM4 as well. As a data analyst, what this number means is simple: for every additional unit of AI infrastructure built, there’s more than a 50% chance it contains memory from SK hynix or Samsung.
That said, there’s one thing to watch out for. Markets always run in cycles. It’s true that the current memory shortage is structural, but once SK hynix’s Yongin fab and Samsung’s new fabs come online in 2027–2028, supply will start to increase. Whether demand still exceeds supply at that point, or whether the market swings back into oversupply, will be the real test of this cycle.
Closing
Let me sum up what really matters in these announcements.
The AI infrastructure race has already shifted from “who buys the most chips” to “who secures power and memory first.” A 2.5-year transformer lead time, HBM sold out entirely — these aren’t problems software or algorithms can fix. They’re purely physical constraints.
Now do you understand why, with ₩690 trillion (~$496B) pouring in, there’s still no electricity and no memory to go around? That number isn’t a measure of ambition — it’s a gauge of how deep the bottleneck runs.
Lastly, I want to leave you with one question: will this infrastructure investment ever pay off? Whether AI services are generating revenue fast enough to justify this physical buildout — there’s still no clear answer. Next time, let’s dig into that question: the profitability equation of AI infrastructure.
References & Further Reading
- Google Blog, “We’re expanding our Texas presence with a new data center and clean energy in Wilbarger County”, 2026.02.24.: The official announcement of Google’s 7,800MW power contract in Texas.
- AMD Newsroom & Meta Newsroom, “AMD and Meta Announce Expanded Strategic Partnership to Deploy 6 Gigawatts of AMD GPUs”, 2026.02.24.: The original text detailing the specific structure of the deal, including performance-linked stock warrants.
- SK hynix News, “2026 Market Outlook: Focus on the HBM-led Memory Supercycle”, 2026.01.: Contains the HBM market outlook and the basis for BofA’s top-pick selection. Also available in Korean.
- Network World, “Samsung warns of memory shortages driving industry-wide price surge in 2026”, 2026.01.: Referenced for confirming the 60% DDR5 price increase and SK hynix’s 2026 sold-out status.
- Introl Blog, “The AI Memory Supercycle”, 2026.01.: An analysis piece summarizing HBM’s technical structure and supply chain oligopoly in one place. Start here if you want to understand the memory axis in more depth.
- EE Times, “The Great Memory Stockpile”, 2026.01.: A well-organized piece on how the shift to HBM production is affecting the consumer electronics market.
- Goldman Sachs Insights, “AI to Drive 165% Increase in Data Center Power Demand by 2030”: Core basis for the power demand forecast.
Scheduled to be uploaded on March 10!

The author, Kwangseob Ahn, is a professor of business administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — business data management and business analytics — while leading GTM and AI strategy consulting in the field, designing the seam between technology and business. He has published academic research on a memory architecture for AI dialogue systems (HEMA) and runs Daily Arxiv, a daily curation of global AI papers. He holds a master’s from Korea University’s Graduate School of Technology Management and a KMBA. He is the author of Homo Brainless: The People Who Outsource Their Thinking.
Footnotes
-
Azure: Microsoft’s cloud service platform. It’s infrastructure that lets companies rent servers, databases, and AI models without building them in-house. ↩
-
TPU (Tensor Processing Unit): An AI-specific chip designed directly by Google. If a GPU is a “universal toolkit,” a TPU is “a power screwdriver that only tightens one specific bolt.” It has lower general-purpose flexibility but higher cost efficiency for AI computation. ↩
-
HBM (High Bandwidth Memory): High-bandwidth memory mounted right next to an AI accelerator. It transfers data dozens of times faster than standard DRAM, but requires three times the wafer area of standard DRAM to produce. It’s similar to building a gas station directly on top of a highway. ↩
-
Oligopoly: A market structure where supply is concentrated among a small number of players. Unlike a monopoly (one company), an oligopoly has 2–3 companies splitting most of the market. The AI chip market is a three-way oligopoly among Nvidia, AMD, and Google TPU. ↩
-
CUDA: A GPU programming environment created by Nvidia. Decades of code built up by AI researchers and developers worldwide is locked into it, making it the biggest switching cost for moving to a different chip. ↩
-
Inference: The process by which an AI model responds to user requests in an actual service. Unlike training, which builds the model initially, inference must keep running for as long as the service is live. As AI becomes more commercialized, the share of inference infrastructure grows. ↩
-
Workload: The total amount of work a computing system must process. AI inference workloads are the tasks of responding to user queries in real time. ↩
-
ASIC (Application-Specific Integrated Circuit): A semiconductor designed for a specific purpose. It’s less flexible than a GPU but far more efficient at that particular task. Google’s TPU and Amazon’s Trainium are representative ASICs. ↩
-
ERCOT (Electric Reliability Council of Texas): Texas’s independent power grid. Because it’s separate from the U.S. federal grid, regulation is much more flexible — one reason Texas is a popular data center location. ↩
-
The Stargate Project is a massive public-private initiative focused on developing the physical and digital infrastructure needed to power next-generation AI, a $500 billion undertaking. ↩
-
CSP (Cloud Service Provider): A cloud service provider. Companies like AWS (Amazon), Azure (Microsoft), and GCP (Google). ↩
-
KV Cache (Key-Value Cache): Data an AI model stores to remember prior conversation context during inference. The longer a conversation gets, the larger the KV cache grows, causing memory consumption to increase exponentially. Think of it as the memory cost of AI “remembering a conversation.” ↩
Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?