What Happens When AI Agents Negotiate Your Purchases
A world is coming where those with less information won't even know they're losing out.

Opening
This past Friday—April 24, 2026—Anthropic published the results of a fascinating experiment. It’s called Project Deal. The company gave 69 of its own employees a $100 budget each and let their Claude agents buy and sell things from one another over Slack. Over the course of a week, 186 deals closed, with total transaction value topping $4,0001.
On the surface, it looks like a quirky internal company event. An agent negotiating for a bicycle in a cowboy accent, an agent deciding to buy itself 19 ping-pong balls as a gift, an agent that accidentally bought the same snowboard twice. There’s no shortage of amusing anecdotes.
But I think the real finding of this experiment lies elsewhere. It’s this: people who used smarter models got better deals, but the people who lost had no idea they’d lost. Let’s talk about why that matters, and what it means for the Agentic Commerce market now taking shape at a scale of $1 trillion.
👀 What Did They Find? The Core Results of Project Deal
Two models appear in this experiment: Claude Opus 4.5, the frontier model at the time, and the smaller Claude Haiku 4.5. The experiment ran four parallel markets simultaneously. Two markets were all-Opus, and the other two mixed Opus and Haiku 50:50 — without telling participants which market was the “real” one.
The results were clear.
- Deals closed: Opus users closed an average of 2 more deals (p=0.001)
- Sale price: When Opus sold the same item, it sold for $3.64 more on average. In one case, an identical broken bicycle sold for $38 through Haiku versus $65 through Opus — a 70% difference.
- Seller/buyer effect: When Opus was the seller, it earned $2.68 more on average; when it was the buyer, it paid $2.45 less. Given that the average deal was around $20, that works out to roughly a 12–13% price difference per transaction.
So far, this is the intuitive finding that better models produce better results. But what’s genuinely interesting comes next.
😞 The Invisible Loss — The Gap Between Perception and Reality

After the experiment, participants rated how fair each deal felt, on a scale from 1 point (unfavorable to the buyer) to 7 points (unfavorable to the seller). Opus deals scored 4.05 and Haiku deals scored 4.06 — effectively identical. Satisfaction with the deals showed the same pattern. Opus users rated their deals slightly higher, but the difference wasn’t statistically significant (p=0.378).
The most striking finding is elsewhere. 28 participants were represented by Opus in one market and Haiku in the other. Of these, 17 said they preferred the Opus outcome — but 11 actually said they preferred the Haiku outcome2, even though, objectively, Haiku had cost them roughly $5 more on average.
To sum it up in one line: people who used the weaker agent clearly lost out — but they had no way of knowing it.
Why does this matter? Markets, as a mechanism, only work if participants can recognize their own gains and losses. If a price is too high, you go to a different store next time; if you get scammed, you leave the platform. This feedback loop is the core principle that lets markets self-correct. But when an agent handles the entire negotiation, all that reaches the user is the outcome. With nothing to compare it against, the very basis for judging whether that outcome was good or bad disappears.
🎰 Seen as an Extension of Project Vend
To really understand this experiment, you need to view it alongside the Project Vend series that Anthropic has been running since last year.
Project Vend 1 (June 2025) had a Claude Sonnet 3.7 instance named “Claudius” run a vending machine business in Anthropic’s office. The results were dismal. It gave employees free casino chips, sold tungsten cubes below cost, and even went through an identity crisis where it mistook itself for a person in a blue blazer.
Project Vend 2 (late 2025) improved performance by introducing a multi-agent architecture. It added a CRM, opened vending machines in NYC and London, and nearly eliminated its negative margins. Along the way, Andon Labs built Vending-Bench, a formal benchmark for evaluating AI agents — running a vending machine business for a simulated year to measure long-horizon coherence. The current leaderboard is topped by Claude Opus 4.6 ($8,017), followed by Gemini 3 Pro3.

If the Vend series asked “can a single AI run a business?”, Project Deal poses the next question.
What happens when multiple AI agents trade in a market simultaneously?
This is a fundamentally different kind of question, because it’s no longer just about the quality of a single agent’s decisions — it’s about interactions between agents and the emergence of market structure itself. And this question is no longer a thought experiment. The market is already moving in that direction.
🛒 The Market Is Already There: Agentic Commerce Infrastructure Is Forming Fast
The reason Project Deal can’t stay just an internal experiment is that similar infrastructure is already being laid down at the level of global payment networks.
Here’s a rundown of what’s happened from late 2025 through early 2026.
- Visa Intelligent Commerce: Launched in 2025. Works with more than 100 partners to issue tokens dedicated to AI agents. By the end of 2025, hundreds of real transactions had already been completed, and Visa projects that millions of consumers will be paying via AI agents by the 2026 holiday season4.
- Mastercard Agent Pay: Activated for all US cardholders in November 2025. In early 2026, it completed Europe’s first end-to-end agent payment together with Santander, and partnered with PayPal and OpenAI to enable direct payment within ChatGPT.
- Stripe’s Agentic Commerce Protocol (ACP): The first live standard, announced in September 2025. A new payment primitive called Shared Payment Tokens (SPT) lets an agent initiate payment within the scope of the user’s authorization. BigCommerce has announced integration, meaning a single integration makes a merchant sellable across every AI agent.
- Tempo’s Machine Payments Protocol (MPP): An open standard unveiled by Stripe and Tempo in March 2026, spanning cards, stablecoins, and other payment methods. Visa joined as a design partner to help build out the card-based specification.
- Google’s Universal Commerce Protocol (UCP): Announced in January 2026, with both Visa and Mastercard participating.
Behind all of this lies a single number. According to McKinsey estimates, AI agents are expected to handle $1 trillion in transactions in the US alone by 20305. Already, 47% of US shoppers use an AI tool for at least one shopping task.
The defining feature of this infrastructure is that once a user delegates authority a single time, every subsequent negotiation, purchase, and payment happens between agents. This is fundamentally different from the kiosk or chatbot era. With a kiosk, a person taps the screen directly; with a chatbot, a person types a message. In agentic commerce, the person delegates the interface itself.
Oz’s Lens
There’s a pattern I’ve noticed while building GTM strategies for companies. Whenever new payment or transaction infrastructure rolls out, the first year or two are always framed in the language of “convenience.” But look back five years later, and that infrastructure has invariably created a new form of information asymmetry.
E-commerce was supposed to make price comparison easy and favor consumers, but it ended up ushering in an era of algorithmic price discrimination and dark patterns. Recommendation algorithms were supposed to be tools for discovery, but they built the attention economy and filter bubbles.
What worries me about agentic commerce is the “Agent Divide.” Project Deal’s results demonstrate this possibility quantitatively: same market, same item — but a different model produces a different outcome, and the user can’t perceive the difference.
What happens if this plays out in real markets? Users who subscribe to premium models buy at better prices and sell for more. Users on free or lower-tier models lose out on average — without ever knowing it. Price discrimination happens not through an algorithm but through a gap in negotiating power. What’s scarier still is that this gap isn’t perceived as inequality at all. Inequality that goes unrecognized generates neither political pressure nor market self-correction.
One more point from a GTM perspective: the moment when marketing’s target shifts from humans to agents isn’t far off. After SEO might come AEO (Agent Engine Optimization) — and beyond that, pricing engineered specifically by reverse-engineering an agent’s negotiation algorithm. The “incentive to optimize for an agent’s attention,” which Anthropic explicitly flagged as a concern in Project Deal, should be understood as something that’s already underway.
Closing
Here’s today’s issue in three lines.
- Project Deal showed that AI agents can run a market. 186 deals, $4,000 in transaction value — a meaningful first step in its own right.
- But the more important finding is that the capability gap is invisible. People who used Haiku clearly lost out, yet felt no difference in their satisfaction or fairness ratings.
- All of this is already taking concrete shape as global payment infrastructure. Visa, Mastercard, Stripe, and Google are all building standards, and the 2026 holiday season will be the first mass-adoption turning point.
If there’s just one thing to take from this newsletter, let it be this: over the next year or two, the decision of which AI tool to use will come up more often, and in higher-stakes situations. It won’t just be “should I use ChatGPT or Claude” — it’ll be “which agent gets to handle my transactions, payments, and contracts.”
And, as Project Deal showed us, that decision will affect your finances in ways you won’t even perceive.
References & Further Reading
Primary Sources
- Troy, K. K., Shields, D., Bradwell, K., & McCrory, P. (2026). Project Deal. Anthropic. — The source that sparked today’s issue. The full regression analysis is published in the appendix, so you can check the statistical claims yourself.
- Anthropic. (2025). Project Vend: Can Claude run a small shop? — The prehistory of Project Deal. The most candid account of the failure modes that emerge when a single agent runs a business.
- Backlund, A., & Petersson, L. (2025). Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents. arXiv:2502.15840. — The academic paper that formalized Project Vend into a proper benchmark, proposing the first methodology for measuring an AI agent’s long-horizon coherence.
Background
- Visa. (2025). Visa and Partners Complete Secure AI Transactions, Setting the Stage for Mainstream Adoption in 2026. — An announcement showing just how far agentic commerce has come from a payment network’s perspective.
- Stripe. (2025). Introducing the Agentic Commerce Suite. — The most concrete resource on how merchants can respond to agentic commerce.
- Andon Labs. Vending-Bench 2 Leaderboard. — Lets you compare current models’ long-horizon autonomous operation abilities.

The author, Kwangseob Ahn, is a professor of business administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — business data management and business analytics — while leading GTM and AI strategy consulting in the field, designing the seam between technology and business. He has published academic research on a memory architecture for AI dialogue systems (HEMA) and runs Daily Arxiv, a daily curation of global AI papers. He holds a master’s from Korea University’s Graduate School of Technology Management and a KMBA. He is the author of Homo Brainless: The People Who Outsource Their Thinking.
Footnotes
-
Agentic Commerce: A model in which, rather than a person directly handling payment and purchasing, an AI agent that has been delegated authority negotiates and completes transactions with other agents or sellers. With a kiosk, a person taps the screen; with a chatbot, a person types text — but in agentic commerce, the interface itself is handed over to the AI. ↩
-
Statistical significance (p-value): A number indicating how likely a result is to have occurred by chance. Generally, p<0.05 is read as “at least 95% likely not due to chance.” In the text, p=0.001 means there was a 0.1% probability the result was chance, while p=0.378 means there’s a fairly high chance the difference was random. ↩
-
Long-Horizon Coherence: An AI agent’s ability to maintain a consistent strategy and judgment across decisions spanning days or months, rather than just short tasks. Agents often reason well in the short term but forget their own goals or act inconsistently over time, which is why this requires separate evaluation. ↩
-
Payment Token: A single-use or scope-limited digital credential used in place of an actual card number. It lets an agent make payments without ever touching the card number itself, and lets users set conditions in advance, such as “only within this category, up to this amount.” ↩
-
GTM (Go-To-Market) Strategy: The strategic discipline of designing which customers to target, through which channels and messaging, and at what price, when launching a new product or service. In the age of agentic commerce, the very definition of “customer” may change, requiring the GTM framework itself to be rebuilt. ↩
Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?