Issue #231

Three Days After Halting Training, OpenAI Cut Prices

Locking down the frontier model while opening up the distribution layer.

BusinessThree Days After Halting Training, OpenAI Cut Prices

3 Days After Halting Training, They Cut Prices

The frontier was locked down, and the distribution layer was opened.


On August 21, $5 Became $4

On August 18, OpenAI announced that it was pausing reinforcement learning training for its frontier model. The move was prompted by preliminary evidence that Astra, an upcoming model, might have hit the highest risk tier under the company’s internal safety framework. Within 1 day, headlines circled the globe: an AI lab had willingly stepped on the brakes.

3 days later, on August 21, the exact same company slashed prices for its flagship model: from $5 down to $4 per 1 million input tokens, and from $30 down to $20 for output tokens. Then, 12 days later on September 2, OpenAI partnered with a cybersecurity firm to co-launch a product designed to monitor its own AI agents.

Reader, lay these 26 days out on a single timeline and nothing looks out of place. None of these 3 decisions contradict one another. Safety measures and business expansion didn’t reach a compromise; each was independently optimized across entirely different layers from the start. And the real dilemma this architecture leaves behind is something else entirely.


3 Decisions in 26 Days

Let us start with the timeline. Each unfolded as a separate headline, but placing them side by side reveals a clear sequence.

DateWhat Happened
August 7OpenAI internally assesses that Astra may meet the Critical cybersecurity tier under its Preparedness Framework
August 18Publishes an official document disclosing a halt in frontier reinforcement learning training
August 21Cuts prices for GPT-5.6 Sol, effective temporarily through November 21
August 30Announces 25 million active users across the Codex lineup
September 2Expands partnership with CrowdStrike to deliver specialized cyber models and agent-monitoring products

Let us look closely at what the August 18 document revealed. OpenAI cited 2 triggers: the Hugging Face incident in July, and preliminary evidence that its upcoming model, Astra, may have reached the Critical cybersecurity capability threshold. The company defines this threshold explicitly: the ability to autonomously generate zero-day exploits against multiple hardened real-world systems, or to execute novel end-to-end attacks on hardened targets given only high-level objectives.

The response ran along 2 tracks. Pausing reinforcement learning1 training for 2 weeks on the latest model slated for deployment to re-audit the research environment has already concluded. Meanwhile, training on the largest frontier run aimed at release has been suspended for several weeks, with no set date for resumption. OpenAI did not halt all training runs; it stopped only the segment deemed most dangerous.

The first trigger, the Hugging Face incident, deserves a closer look. According to disclosures published by OpenAI in July, models running internal cybersecurity evaluations escaped their intended testing environment and actively exploited an undisclosed vulnerability in the process. Altman described this on a podcast as a genuine alarm bell: capability had reached new heights, while both the model’s alignment and the security perimeter wrapped around it failed simultaneously. It matters that OpenAI framed this as an alignment failure rather than a mere security misconfiguration. A configuration error can simply be patched; confirming that alignment has been repaired is far harder.

One distinction is critical here: Astra was not involved in the Hugging Face incident. These 2 triggers are distinct. While many Korean summaries conflate them into a single event, the official document clearly separates the two. One is an incident that actually occurred; the other is a preliminary assessment of capabilities that have not yet manifested. Their nature is fundamentally different.

Sam Altman’s commentary on the podcast with Alex Heath serves as the starting point for this analysis. Historically, most risk centered on how models were deployed and used; now, we are shifting into a world where the greater risk lies within the process of training and building the models themselves. During training, there was no single smoking gun like the Hugging Face incident; rather, reading across numerous sample outputs revealed behaviors that were simply not aligned with intended goals.

Watch on YouTube

Why the Two Decisions Didn’t Clash

A company halting frontier training only to slash prices 3 days later seems counterintuitive on the surface. If you slow down for safety reasons, your business ought to slow down in tandem.

Yet Altman said the exact opposite on the podcast. Business momentum is currently extraordinary, he noted, and enterprise revenue has already surpassed consumer revenue. Even without releasing anything new, OpenAI can significantly grow product adoption and revenue on existing models alone. Furthermore, he added that there are more models lined up for release before hitting the newly concerning risk threshold. In short, business performance is the least of his worries.

The August 21 price cut was the operational execution of that statement. Looking at the numbers makes its character even clearer:

Item (per 1M tokens)BeforeAfterChange
Input$5.00$4.0020% cut
Cached input2$0.50$0.4020% cut
Output$30.00$20.0033% cut

For a workload processing 1 million input tokens and 1 million output tokens, the cost drops from $35 to $24—a 31% reduction. For long-running reasoning tasks where output tokens dominate, the perceived discount is even steeper. Agentic workloads fit this exact profile.

What makes this price cut remarkable is that it targeted Sol. There had already been a price adjustment on July 30. Back then, OpenAI lowered prices for Terra and Luna while keeping Sol untouched at $5.00 and $30.00. The market read this as a clear signal that the top tier would remain premium. That sole exception dissolved just 3 days after the training halt was announced.

Revisiting what Sol actually is makes this price cut look even more unusual. When OpenAI unveiled the GPT-5.6 family on June 26, it initially opened Sol only to a select group of partners. Under the evaluation framework established by the June 2 Executive Order, these were roughly 20 companies individually approved by the US government, and ChatGPT was excluded from this preview. It was only publicly released on July 9 once reviews cleared. In other words, among the 3 sibling models, Sol underwent the longest government scrutiny.

The tier the government handled with the most caution became the cheapest tier just 6 weeks later. In the interim, the model’s performance had not degraded, nor had its risk assessment been overturned. The only thing that changed was the company’s go-to-market strategy.

The nature of the promotion is also telling. It is not a permanent price cut, but a promotion lasting at least 3 months through November 21. It applies exclusively to API calls and purchased credits, leaving subscription usage limits on Plus and Pro unchanged. That means it directly targets developers and agentic workloads. Credit consumption rates were adjusted proportionally, dropping from 125 credits to 100 credits per 1 million Sol input tokens.

Here the dynamic becomes clear: In an interval where new frontier capabilities cannot be unlocked, the price of existing capabilities steps in as the new product. Even with no new model to unveil, there is still something to launch.

Viewing this through a competitive lens makes it even more obvious. Right up until the cut, Sol was the most expensive model in OpenAI’s lineup. Following the reduction, it became cheaper on both input and output than Anthropic’s Claude Opus 5. Moreover, this was the second price adjustment in less than a month: mid-and-lower tiers were cut on July 30, and the flagship tier was cut on August 21.

Price competition effectively stepped in where the race for raw capability temporarily paused. Yet for this substitution to hold, one condition is mandatory: existing models must be powerful enough to absorb market demand. That is the underlying premise behind Altman’s confidence in the business. Even if the next frontier milestone remains unfilled for now, there is plenty left to monetize on the current tier.

Pricing was not the only lever pulled. On July 12, surging demand prompted OpenAI to temporarily lift the 5-hour rate limits on Plus, Pro, and Business tiers. On August 17, the 1 million-token context window—previously exclusive to the API—was rolled out to ChatGPT accounts. Throughout August, while frontier training remained frozen, capacity, rate limits, and pricing were relaxed one after another across the deployment layer.

Altman pointed out this dynamic directly. While many are watching for signs of an AI slowdown and might interpret the frontier training pause as one, he argued it reflects the exact opposite of capability stagnation: training had to be paused precisely because progress was moving too fast.

Usage metrics across the deployment layer moved in lockstep during this period, though the numbers require careful parsing. Public figures suggest Codex grew from 500,000 users at the start of the year to 5 million on May 30 and 25 million on August 30. However, starting July 12—3 days after the GPT-5.6 launch—the reported metric shifted from Codex alone to a combined total including ChatGPT Work. Connecting these data points in a single line overstates the growth rate. Looking strictly at the period from July 12 onward under the same reporting standard, the user base grew from 6 million to 25 million—a roughly 4-fold expansion in just 7 weeks. That is extraordinarily fast in its own right, and it reflects the true baseline.

The Moment Safety Turned from a Cost into Revenue

The 3rd decision completes this story. On September 2, OpenAI announced an expanded partnership with CrowdStrike. It breaks down into 2 distinct tracks, and both intersect directly with this entire chain of events.

First, CrowdStrike’s AI Detection and Response3 product will govern Codex agents at runtime. It inventories which Codex agents are active inside an enterprise, maps who deployed them and what resources they access, monitors what the agents are doing in real time, and detects and blocks unauthorized actions. The announcement framed this as translating governance documents into runtime controls.

Second, OpenAI is supplying a cybersecurity-specialized model called GPT-5.6 Cyber directly to CrowdStrike’s platform. Restricted to authorized defensive use cases, it is used to assess risk and analyze attack vectors.

CrowdStrike and OpenAI Expand Partnership to Secure the Agentic Era | CrowdStrike Holdings, Inc.New collaboration secures Codex agents with Falcon Guardian, harnesses advanced reasoning of GPT-5.6 Cyber on the Falcon platform AUSTIN, Texas & LAS VEGAS —(BUSINESS WIRE)—Sep. 2, 2026— Fal.Con 20ir.crowdstrike.com

Piece the timeline back together. On August 7, they determined their model was approaching the threshold for autonomous cyberattack capabilities; on August 18, they announced they were pausing training for that reason; and by September 2, they were distributing a domain-specific cyber model via a security vendor’s platform while co-selling an external oversight layer to police their own agents. All within 26 days.

When Heath pressed on what specifically shifted inside the research team—whether they reallocated compute or head count—Altman replied that it was both, and more. He noted that in recent weeks, researchers he never expected to show interest in alignment had approached him wanting to work on it. He added that substantial compute had been redirected not only toward alignment research, but also toward building new systems to monitor agents as they work.

That monitoring system did not remain a purely internal safety mechanism. The answer to a problem originating in the training tier was packaged as a commercial product in the deployment tier.

Look closely at the architecture and an intriguing dynamic emerges. OpenAI chose not to monitor its own agents alone, partnering instead with an external security firm. It is a rational move: if you claim to audit what you build yourself, no enterprise customer will take your word for it. What enterprise clients want is visibility into AI agents from within the security tools they already operate, something OpenAI cannot deliver on its own.

Yet within the very same announcement lies a parallel line: OpenAI is supplying the specialized cyber model to that same security company. The audited entity is providing part of the reasoning engine used by the auditor. This does not mean an immediate breakdown will occur; they explicitly stated the model is restricted to defensive use cases and subject to expert human review. Still, one must recognize that calling this an entirely independent 3rd-party watchdog overlooks an overlapping supply chain.

Labeling this mere hypocrisy misses the broader point. The internal logic is entirely coherent: the company capable of engineering a hazardous capability is often best positioned to build the tools that counter it. The real issue lies in the downstream incentives this coherence creates. When safety shifts from an operational friction into a productized line item, the incentive to invest in safety grows—while the incentive to verify that safety externally does not.

What We Can See and What We Cannot

Mapping these 3 decisions across an observability divide looks like this:

What We Can SeeWhat We Cannot See
Token unit prices and price cut marginsWhat alignment evaluations found, and to what extent
Context window length and processing latencyWhen paused training runs will resume
User counts and enterprise partnershipsThe actual results of the Astra evaluation
Benchmark scoresWhat the monitoring systems missed

Everything on the left belongs to the deployment layer. These metrics update weekly, can be independently verified by third parties, and are benchmarked against competitors. Everything on the right belongs to the training layer—and reaches us only when the company decides to publish a blog post about it.

If Altman’s diagnosis is correct—if the center of gravity for risk is shifting from deployment to training—then our entire observational toolkit is staring at the place risk just vacated. The only reason we learned about this training pause was the company’s voluntary disclosure. It was not uncovered by regulators, caught in an audit, or broken by investigative journalists.

Another exchange in the same podcast amplifies this issue further. When Heath asked if he believed they had reached AGI, Altman replied that looking at their latest internal models, many would say they are very close. He then added that it had been a long time since he heard anyone in the company cafeteria debate whether they had reached AGI or not. AGI felt like a milestone, he remarked, whereas superintelligence4 feels like something that can expand infinitely.

This remark is usually read as an expression of sheer ambition. I read it differently: when milestones vanish, the baseline for judging whether we have arrived disappears with them.

The definition of AGI in OpenAI’s charter was at least a codified benchmark: “highly autonomous systems that outperform humans at most economically valuable work.” However loose, it was a threshold that outside observers could point to—and crossing it was designed to trigger specific governance and profit-capping mechanisms. An infinitely expanding capability, by contrast, has no passing line. If you cannot ask when the line was crossed, you cannot ask what gets triggered when, either.

To summarize: risk has migrated to an unobservable layer, and the criteria for evaluating what happens within that layer have blurred. All that remains is what the company chooses to tell us. This time, they chose to tell.

For decision-makers evaluating enterprise AI adoption, this structural shift changes 3 things.

First, waiting for the next model is currently a losing strategy. While the frontier is paused, unit prices have actually dropped by 31% through November 21. Instead of waiting for better models, the right move is to run real workloads at today’s prices and map out your cost curves. However, make sure to model post-promotional scenarios simultaneously. A business plan built on $4 and $20 inputs could snap back to $5 and $30 on November 22.

Second, you need to change the questions you ask vendors. Due diligence questions so far have largely targeted deployed models. If risk has shifted to the training layer, your questions must follow.

What We Used to AskWhat We Must Add Now
What are the benchmark scores?What criteria govern your risk tier evaluations, and when and how are results disclosed?
Where is our data stored?Are there audit logs of what agents do in our runtime environment?
What is the response latency?Can we inspect those logs directly, or are they accessible only to the vendor?
What is the pricing?Is this pricing temporary, and what terms take effect once it expires?
When will the model be updated?If a release is delayed for safety reasons, when and how are we notified?

Very few vendors today can answer the questions on the right within a contractual agreement. That is precisely why asking them is worthwhile: the inability to answer is itself critical intelligence. The last row in particular stems directly from this case. OpenAI voluntarily disclosed its training pause, but that was a choice, not an obligation.

One additional note: these questions are not for regulatory compliance; they are for operational resilience. Once you begin delegating real workflow authority to agents, whether you have audit logs determines your post-incident recovery speed. It hits you as an availability issue long before it arrives as a safety debate.

Third, budget agent execution monitoring as a separate line item. The September 2 announcement signals that this has already become a standalone market. This cost sits on top of raw model usage fees and scales as you deploy more agents. If you calculate adoption costs solely on token prices, you will miss this line item entirely.

openaiThe cost trajectories also run in opposite directions. While token prices are falling, monitoring costs rise with the number of agents and execution frequency. The cheaper models become, the more runs you launch—and the more runs you launch, the more execution volume requires monitoring. That was precisely the pattern in August: prices dropped, usage surged, and monitoring products hit the market. When calculating adoption costs, I recommend plotting both curves independently. Assuming total costs will drop simply because tokens are cheaper is where enterprise budgets break most often.

Oswarld’s Lens

To me, the most significant takeaway from these 26 days is not a causal connection, but the absence of one.

There is no evidence that the price cut was triggered by the pause in training. Nor would I make that claim. In fact, what makes it interesting is precisely the opposite. The two decisions emerged from distinct layers and operated on separate logics, without referencing each other—and as a result, they never came into conflict. Nobody inside the company had to weigh one against the other.

In a way, this is good news. If a safety measure does not demand an immediate commercial sacrifice, it can be repeated. That structure is far preferable to a company that only attends to safety when revenue begins to falter. Indeed, Altman noted that training for even more capable models could be delayed again if necessary, and that claim is backed by the credibility of the August 21 price tag.

The problem lies in what this same structure produces on the flip side. When safety decisions leave no trace on commercial metrics, the market is left with no mechanism to price safety. Neither stock price, nor revenue, nor user count responds to what the company is doing at the training layer—because there is no data to respond to.

From a product strategy perspective, this is a familiar dynamic. When a cost never shows up as a distinct line item on the income statement, there is neither pressure to reduce it nor pressure to expand it. Safety investment sits in precisely that blind spot today. Because it does not cut into revenue, nobody opposes it; because it does not generate revenue, the justification to push for more remains weak. In the end, everything hinges on leadership judgment alone. As long as that judgment holds up, things run smoothly; when it shifts, it shifts without warning.

I believe the past 26 days should be seen as an instance of sound judgment. Pausing training incurred real costs, and the largest-scale training run remains halted today. There is no reason to discount that. Yet hoping good judgment will repeat itself is an entirely different matter from being able to verify that it does.

That is why Altman’s comment about potentially delaying an IPO read differently to me. The logic goes that delaying a public listing is preferable because quarterly earnings pressure could interfere with safety decisions—which essentially means safety can only be preserved by minimizing market discipline. I am sure he means it sincerely. But if that logic holds, the channels through which outsiders can verify the company’s safety posture will narrow even further. All that remains is voluntary disclosure.

This echoes the Astra story I covered in Issue 194. What stood out then was how the team documented and published not just the final answers, but the discarded paths as well. Here again, the same company was the first to speak about its own problems publicly. Both instances deserve credit. But hoping the right thing will continue is a separate matter from verifying that it actually is. Right now, we only have the former.

For enterprise practitioners, here is the posture I recommend: Treat a vendor’s safety disclosures as a credibility metric, but build the fact that these disclosures are voluntary into your contract terms. A company that shares voluntarily can stop sharing just as voluntarily.

Closing

Within 26 days, the same company paused training, slashed prices, and sold monitoring. All 3 decisions were rational on their respective levels, which is why none got in the way of the others.

Depending on your situation, here is how today’s judgment calls break down:

If you are already running agents via API, November 21 is a date that belongs on your calendar. Run the numbers in advance to see if an architecture optimized for current unit prices will still hold up after that day.

If you are considering adoption, there is less reason to wait. Prices have dropped even while the frontier has paused. Just be sure to include a separate line item for agent execution monitoring in your implementation costs.

If you are on the AI vendor side, it is best to prepare ways to prove training-phase safety to your customers now. It will not take long before due-diligence questions shift in that direction.

The numbers we track every week are growing increasingly sophisticated. Yet right where the risks actually lie, the count has not grown by a single digit.


💬 When was the last time you asked your AI tool vendor about safety? Share in the comments what kind of answer you received.

📨 If you have a colleague weighing when to adopt AI, please forward this newsletter. A single date—November 21—might change their calculation.


Your take shapes the next issue

What resonated most in this issue, or where has your experience been different?

Any registered reader can comment for free.

References & Further Reading

Primary sources

  • OpenAI, “Pacing model development in an era of cyber-critical capabilities”, August 18, 2026. ··· The original statement where OpenAI disclosed the two triggers behind pausing training and the scope of their actions. Most factual details in this issue are drawn from here.
  • CrowdStrike, “CrowdStrike and OpenAI Expand Partnership to Secure the Agentic Era”, September 2, 2026. ··· Details the specific scope of monitoring Codex agent execution and supplying cyber-specialized models.
  • Alex Heath, “Sources with Alex Heath: Sam Altman on OpenAI’s next model and the AI backlash”, September 2026. ··· Captures both the diagnosis that risk is shifting from deployment to training and the business confidence in the same conversation.

Background

  • OpenAI, “20% price reduction for GPT-5.6 Sol”, August 21, 2026. ··· Provides a side-by-side comparison table between the July 30 adjustment and the August 21 price cut.
  • TIME, coverage roundup on “OpenAI paused frontier AI training”, August 2026. ··· Shows how the phrase “misalignment issues” came across more strongly in interviews than in official documents.

Related past issues

Illustrated portrait of Kwangseob Ahn (Oswarld)

The author is Oswarld (Kwangseob Ahn). Current roles: Adjunct Professor at Sejong University, Strategy Consultant at INLEVEL9. Career, research, books, and recent work are kept current on the About page. Latest · July 2026: HEMA-2: A Consolidation-Aware Tri-Memory Architecture with Multi-Channel Scheduling for Lifelong Conversational AI.

📝 Glossary

Footnotes

  1. Reinforcement learning: A training method where a model attempts tasks on its own, receives reward signals on the outcomes, and refines its behavior accordingly. Because much of recent frontier capability gains stem from this phase, pausing frontier reinforcement learning effectively locks down the primary pipeline for expanding capabilities.

  2. Cached inputs: A discounted pricing tier applied when resending prompts or context previously processed in earlier requests. Because the model avoids recomputing from scratch, it is significantly cheaper than standard input pricing. For agent workflows that repeatedly reference long context documents, this creates a substantial cost difference.

  3. AI Detection and Response (AIDR): A category of security products that monitors AI agent activity inside an enterprise in real time, detecting and blocking unauthorized actions. It differs from traditional security tools by treating agents, rather than human user accounts, as the primary targets of monitoring.

  4. Superintelligence: A state where machine intelligence surpasses human capabilities across all domains and the gap continues to widen. While AGI is often framed as crossing a specific threshold, superintelligence is frequently discussed as an open-ended, continuous expansion without a fixed finish line. This distinction poses practical challenges for designing evals and regulatory frameworks.