The Moat Didn't Move. The Buyer Did.
AMD didn't win Advancing AI 2026 with a faster chip. It won by making switching cheap. That's a lesson about your go-to-market, not their silicon.
AMD did not win the AI argument last week with a faster chip. It won by making the cost of switching disappear.
I spent two days at Advancing AI 2026. In the keynote hall, in the workshop rooms, and in the hallway conversations with the builders, investors, founders, and developers who actually ship this stuff. The thing that mattered most had almost nothing to do with gigawatts.
Patrick Moorhead, who has covered AMD for thirty years, titled his field note “Helios Ships, Venice Swings, And The Software Moat Narrows.” That third point is the whole game from the Go-to Market perspective. The moat was never the hardware. It was the porting cost, the years of rewriting and tuning it took to move a workload off the incumbent.
That cost is collapsing. And when a switching cost collapses, buyers get choice. When buyers get choice, orchestration decides the winner, not the spec sheet. That is a GTM story. It’s your story.
The moat you're defending is made of porting cost
Here’s the moment that stopped me. On stage, Philippe Tillet of OpenAI, the man who created Triton, described AI agents now writing optimized GPU kernels down to instruction scheduling in assembly, against an open compiler stack. Work he said would have been impossible in a closed one.
Vamsi Boppana, who runs AI at AMD, put a number on it. His team ran 14,000 models through an AI optimizer called Hyperloom in a single pass, something no human team could do by hand. And Anthropic’s story, retold all week, was one engineer wiring Claude to a rack over a weekend and coming back to a working performance graph.
Read that as an operator, not an engineer. The most defensible moat in tech, the reason customers stayed put for a decade, was integration labor. AI just turned integration labor into a weekend job.
So ask the uncomfortable question. What is your moat made of? If the answer is “the work it would take a customer to leave,” you’re defending a porting-cost moat. And porting cost is collapsing in your category too. It’s only a question of which quarter.
Tokenomics is the new pipeline math
Every serious room was really talking about one thing: the cost per unit of intelligence. AMD calls its edge “tokens per dollar.” Dan McNamara said AMD’s own IT team cut token costs 43% by routing each request to the right model on the right compute instead of sending everything to the frontier.
Then AT&T’s CTO Jeremy Legg made it real. Over a trillion tokens a month. More than a hundred GenAI models in production. Cache-aware routing cutting some AI costs by as much as 90%. He used the word I’d been waiting for a Fortune 50 operator to say out loud: moneyballing. Moneyballing token costs across a portfolio of open and closed models to hit the same outcome for a fraction of the spend. That’s not an infrastructure story. It’s a revenue-efficiency story in an infrastructure costume.
Here’s the translation. Tokens are becoming a cost of goods sold. The teams that measure, route, and compound their token economics will out-margin the teams that don’t, exactly like the teams that measured pipeline conversion buried the ones that guessed.
Tokenomics is the new pipeline math. If your RevOps function can’t see cost-per-outcome across your AI motion yet, that’s where your margin is leaking.
The distributed stack is an orchestration problem
Watch how AMD framed the day and you’ll see your own operating model staring back. The pitch wasn’t one god-box. It was intelligence distributed across the data center, the enterprise server, the desk, and now the robot. Frontier model in the cloud, open model on-prem, a small model on a device beside you.
Jack Huynh showed personal AI running 300-billion-parameter models on a box that sits on a desk. In the workshops, I watched a small vision-language-action model drive a robot arm on a Ryzen board, and an open engine called Atom serve a coding agent locally with no data center in sight. Cisco’s Jeetu Patel said the line of the conference: “Humans click, but agents swarm.”
Sit with that. The old motion assumed a human at every step. The new one assumes hundreds of agents running around the clock across a distributed stack. That only creates value if something orchestrates it. Routes it. Governs it. Prices it.
This is the AI Orchestration thesis, and the four layers map cleanly. Signal is the router deciding what each request actually needs. Research is the context each agent pulls to make a good decision instead of a generic one. Outreach is the swarm doing the work at machine scale. RevOps is the control plane: the tokenomics governance, the observability, the guardrails. The thing Cisco literally shipped as “Cloud Control” so a leader can quarantine a runaway agent before it torches the budget.
This isn't theoretical, and it isn't only a hyperscaler game. On Claire Vo's How I AI, a solo builder named Alex Finn walked through the exact model running on his own desk. Five machines, five agents, a fleet dashboard he built himself, routing each task to the right local model and running the whole thing 24/7 without babysitting it. Signal, Outreach, and a control plane, operated by one person. And his sharpest line is the tokenomics punchline: unlimited local inference changes the use-case math in a way a $20 cloud subscription never can. One builder already lives in the world AMD spent two days describing.
AMD spent two days proving that in silicon. The lesson for GTM is identical. You don’t have a tooling problem. You have an orchestration problem.
The buyer already voted
Now rewind seven weeks. On June 2 I was at Microsoft Build, watching Satya Nadella lay out the same future from the buyer’s side of the table.
Microsoft didn’t pledge loyalty to a vendor. It bragged about choice. First-party silicon. Partner systems. The first cloud to stand up Nvidia’s Rubin. Deeper work with AMD. And its own Maia accelerator, which Microsoft said delivers 30% better tokens per dollar than the leading GPU, already validated on a frontier model and running Microsoft 365 Copilot. The biggest AI buyer on earth told its developers that silicon is now a portfolio decision, not a marriage.
Same with models. Foundry now lists more than 11,000 of them, and Satya said the model itself is becoming less of the differentiator. When a buyer has 11,000 interchangeable options, the moat isn’t the model. It’s the orchestration around it.
Even Jensen felt it. Beamed in from Taipei, the most powerful incumbent in the industry sold Nvidia not on lock-in but on tokens per dollar per watt, and on making token generation 30 times cheaper than the last generation. When the leader starts selling unit economics instead of switching cost, the moat is already gone. He just hasn’t said so.
That’s the demand signal. Now connect it to the supply side. Seven weeks later, Andrew Feldman and AMD answered that exact buyer. Cerebras took the one axis Nvidia still owned, ultra-low latency, and fused it with AMD’s throughput and memory into a single disaggregated solution. Five times the throughput while keeping the speed. On Cerebras Cloud this year.
That’s not two vendors making a chip. It’s two challengers composing a product to attack the incumbent’s last stronghold, aimed at a buyer who already said out loud it wants choice. June 2 was the buyer voting for a world without a moat. July 23 was the ecosystem shipping the product that world demands. Two months, thesis to purchase order.
The lesson most vendors will miss: in a choice-rich market, the winning play isn’t “do everything myself.” It’s “compose the best combination and let partners vouch for me.” If you’re still running a hero-seller, do-it-all, closed motion, you’re running Nvidia’s strategy in a market that just rewarded AMD’s.
What I took from the rooms
I didn’t go to San Francisco to admire hardware. I went to read the demand.
The builders and founders weren’t asking which chip is fastest. They were asking how to run more intelligence for less, everywhere, without hiring an army. That’s the scale-revenue-without-headcount question, asked by the people building the picks and shovels. The investors were circling the same thing from the other side: where does margin compound, who owns the orchestration layer, and who’s just reselling tokens.




And the developers in the workshops, wrist-deep in ROS 2 and inference engines, were quietly building the routing and control muscle every enterprise is about to need and almost none of them have. That gap is the opportunity. It’s the one I intend to lead into.
The move
Stop buying tools. Start composing the system. The next eighteen months belong to operators who treat AI as an orchestration layer with real unit economics, not a shelf of point solutions with a demo and a logo.
AMD just showed the industry what that looks like in silicon. It made the switch cheap, made the economics legible, and let the ecosystem sell for it. Run the same play in your go-to-market.
So here’s my question. If a customer could leave you as easily as an engineer now moves a workload off CUDA, what would actually make them stay? Reply and tell me. I read every response.











