How You Sell GPU Compute Changes Your Margins
Buying servers, subleasing them, or billing by API token each demand different upfront cash and payback timelines.
BusinessSame GPUs, Different Businesses
There’s more than one way to sell compute to customers who need GPUs. An operator can buy servers outright and rent them out. Or it can lease someone else’s servers and sublease them to customers. Or it can run an AI model on its own servers and charge customers based on how much they use through an API.
Depending on what you acquire and what you actually sell, the upfront investment, the margin left over, and how long it takes to recoup that investment all change. That’s why saying “the GPU business is profitable” doesn’t tell you much on its own — you don’t know which business is being described. Let’s walk through the economics of all three approaches, one at a time.
1. Buy Servers Outright and Rent Them Out
In this model, an operator buys GPU servers, runs them in a data center, and leases them to customers. Customers secure access to a server or GPU capacity and pay by the hour or month.
The operator pays for the hardware upfront. Afterward, rental income covers operating costs like electricity, cooling, and maintenance, and whatever cash remains goes toward recouping the initial investment. Because the hardware costs so much upfront, customers need to keep using the servers for a sufficiently long stretch of time.
What matters most in this model is how cheaply the servers were bought and how consistently they can be rented out. If there are no customers and the servers sit idle, the money already spent on hardware doesn’t come back. And if new GPUs hit the market and drive down rental rates for existing servers, payback can take longer than originally expected.
Locking in long-term rental customers makes revenue easier to predict. But in exchange, the operator has to build out the hardware and operations to deliver on the performance and service duration promised in that contract.
2. Rent servers, then re-lease them
This model involves renting GPU servers owned by another company and then reselling access to customers. The operator handles customer acquisition, setting up the usage environment, and technical support—earning income from the gap between what it charges customers and what it pays the supplier.
For example, if you rent a server for ¥200,000 (~$27,500) a month and charge customers ¥250,000 (~$34,400), the margin is ¥50,000 (~$6,900) a month—20% of revenue. Factor in sales and customer support costs on top of that, and actual profit shrinks further.
Since you’re not spending heavily to buy servers, there’s little pressure to recoup equipment costs over several years. Instead, contract terms become critical. If you’ve committed to a supplier for a year but a customer leaves after just one month, the operator has to absorb the rent for the remaining period. You also need to prepare a security deposit and working capital to cover the gap before customer payments come in.
The first thing to examine in this model is the spread between purchase and resale rates, and the terms of both contracts. In exchange for lower upfront equipment investment, you need to decide how to cover costs during periods when capacity goes unsold.
No Hangul, numbers, or structural issues found in this fragment.
3. Offering the Model via API and Charging by the Token
This is the approach of providing an API so customers can send requests to an AI model and receive results back. Tokens are the units a model breaks text into when processing it. In this business, you charge based on the number of input and output tokens. This kind of model-provision business is called MaaS (Model as a Service).
Server-rental customers run models or programs directly on the resources they’ve secured. API customers, by contrast, just send the requests they need and get back the answers. Behind the scenes, it’s the API provider who manages which servers to use and how to handle requests from many different customers. The provider might use servers it purchased itself, or servers it rents from outside.
Here, revenue isn’t determined by how many GPUs you own alone. What matters together is how many paid requests come in, how many requests the same equipment can process, and how much you charge per token. If you batch requests and process them efficiently, you can generate more revenue from the same servers. Conversely, if you have few customers or price competition drives down unit rates, expected profits become hard to achieve.
Running the model and developing the service also cost money. If you’re developing your own model, training and research costs get added on top. And if you own the servers outright, you also need to recoup the equipment investment through API revenue. So the payback period for an API business varies depending on how you acquired your equipment and how much you invested in the model and service.
How China’s Analysis Shows Different Returns and Payback Periods
A Morgan Stanley analysis cited by Wall Street CN compared these three models. For the direct-purchase scenario, the assumptions were a price of ¥8,000,000 for a server equipped with eight GPUs and monthly rental income of ¥250,000. The API model assumes the operator uses its own servers, with separate assumptions for throughput, the share of paid inference, and token pricing.
Operating margin shows how much is left after subtracting costs from revenue, while ROIC (return on invested capital) shows how much after-tax operating profit is generated relative to capital deployed in the business. Cash payback period is how long it takes to recoup the initial investment through operating cash inflows — this differs from the depreciation period, which is an accounting convention for spreading equipment costs over time.
| Business model | Operating margin (per analysis) | ROIC | Cash payback period |
|---|---|---|---|
| Buy servers outright, then rent out | ~44% | ~13% | ~3.1 years |
| Rent servers, then re-rent them | ~20% | Not calculated (no equipment investment base) | N/A — no equipment cost to recover |
| Provide model API via own servers | ~53% | ~19% | ~2.5 years |
The figures in this table aren’t the actual performance or guaranteed returns of Chinese operators as a whole — they’re a scenario built from the conditions laid out in the article. In particular, the per-server calculations don’t include company-wide R&D spending, personnel costs, and the like. The API model, too, only approaches the table’s results if the assumed paid demand and processing efficiency are actually met.
The takeaway from this comparison isn’t that API is always the better deal. It’s that when the burden of owning equipment, the rent paid to outside parties, and the unit you charge customers for all change, the profit structure changes too — even when you’re working with the same GPUs.
Domestic companies play different roles in this market too
In Korea, VESSL AI and Elice also provide GPU cloud services, while Lablup offers Backend.AI, a platform that allocates and manages GPU resources. They’re all connected in that they supply computing power and usage environments to customers who need GPUs.
Naver Cloud, NHN Cloud, SK Telecom, and KT Cloud also run GPU server and GPUaaS businesses. Megazone Cloud supports the supply and buildout of cloud and AI infrastructure. And there are services that provide model APIs too, like Naver Cloud’s CLOVA Studio.
A single company can run several of these models at once. So rather than judging a company’s revenue structure by its name alone, you need to look at who actually supplies the hardware behind a given product and what exactly the customer is paying for. Equipment supply, leasing, operating software, and model APIs each generate revenue and cost at different points.
Equipment and usage records also get used for loans
How to raise the initial money you need is also a critical part of this business. I covered GPU-collateralized loans in the U.S. and loans based on token usage records in China in my ZDNET Korea column on September 5th.
In the U.S., CoreWeave has raised capital by pledging assets including GPUs as collateral. According to the company’s disclosure, a $2.3 billion loan in 2023 was meant to fund the purchase of GPU servers and other equipment needed to fulfill customer contracts. The relevant subsidiary’s assets and equity were pledged as collateral. It’s a structure where you first secure purchase funds based on equipment and customer contracts.
China has cases of using usage records that demonstrate real business activity in loan screening. The electricity-bill-based lending introduced by the National Data Administration uses small and medium-sized enterprises’ power consumption data in banks’ risk assessments. The token-based lending in Haizhu District, Guangzhou covered in the column feeds operational information—token usage volume, contracts, payment collection history—into loan screening.
Here, usage volume isn’t collateral to be liquidated the way a physical GPU is—it’s material used to judge whether a company is actually doing business and can repay its debt. The significance is that not just companies that own equipment, but also companies that buy and use computing resources, can demonstrate their fundability through their own transaction and usage records.
Once loans become available, the need to raise the full equipment cost with your own money alone decreases. But you still have to cover operating costs from customer payments and repay both principal and interest on the loan. You also need to check whether the timing of recovering the money spent on equipment lines up with the timing you’ve committed to repaying the bank.
Oswarld’s Lens
While building GTM strategies, I’ve made and received plenty of data showing how much profit a single customer, a single store, or a single server generates. The profitability at that individual unit level often looks good, but the company’s overall operating margin doesn’t improve as much as expected even as time passes. That’s because the cost of the people and systems needed to actually run the business was left out of the calculation.
Because of that experience, whenever I look at profitability data, I start by connecting what’s being sold to the customer with what money has to be spent first to generate that revenue.
If it’s a business that leases equipment directly, you need to look at the equipment cost alongside how long the customer uses it. If it’s a business that re-leases capacity, you need to look at both the supplier contract and the customer contract. If it’s an API business, you need to look at paid requests together with processing costs. On top of that, you have to add the people needed for customer support and service operations, plus the cost of capital — only then do you get close to what the company actually keeps.
The preparation needed to scale also differs by business model. You might need to buy more equipment, expand supply contracts, or improve the service so the same equipment can handle more requests. Understanding these differences is what lets you decide where to invest and which customers to pursue.
Closing
To understand the profitability of a GPU business, you first need to separate out the sales model. If you buy the hardware outright and lease it, you have to recoup a large equipment investment. If you rent capacity and re-lease it, you’re managing the spread between rates and contract terms. If you sell through a model API, what matters is paid usage volume and the cost of processing requests.
Once you understand this structure, the numbers behind margins and payback periods start to make sense. Add in the terms on which a company borrows the capital to buy equipment or fund operations, and you can see why companies running what looks like the same AI infrastructure business end up with such different results.
💬 Have you looked into building your own GPU capacity versus renting it? Tell me which factor — equipment cost, contract length, or operating headcount — mattered most in your decision.
📨 If you have a colleague evaluating a GPU or AI cloud business, or weighing the cost of adopting one, please pass this along.
Keep the perspective, not the noise.
We choose one consequential shift and trace what sits beneath it, every other day.
Confirm once to finish subscribing.
Already a subscriber? Sign in to join the conversation
References & Further Reading
- Wall Street CN, introducing Morgan Stanley’s profitability analysis of China’s AI compute business: an article comparing pricing, return, and payback-period scenarios across three business models.
- Kwangseob Ahn, “Bank Loans: In the US and China, GPUs and Tokens Decide — What About Korea?”, ZDNET Korea: my own column on how GPU assets and token usage records are each used in financing.
- CoreWeave’s SEC filing: confirms the purpose and collateral scope of the 2023 loan.
- China’s National Data Administration case on electricity-bill-based lending, Guangzhou Haizhu District’s token-based lending announcement: these explain, respectively, underwriting based on power-usage data and on token/contract data.
- The roles of domestic Korean companies are described based on each company’s official product pages and announcements, linked in the body.

Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?