Margin of Safety #62: Alex Karp’s CNBC Rant and the Battle for Enterprise Sovereignty
Alex Karp’s CNBC tear is 1) quite entertaining and 2) exposed the quiet fury inside corporate IT boards
If you watched Palantir CEO Alex Karp on CNBC recently, you might have initially written it off as another one of his signature, highly animated performance pieces (aka rants). In some ways, it was a calculated and self-serving pitch for Palantir, but his subsequent virality came from highlighting, in a maximally opinionated fashion, a set of major tensions and fears within the AI landscape.
(Source: CNBC)
If you haven’t seen the interview and want a window into the AI zeitgeist, it’s here. Karp focuses his complaint on three points:
1. Nobody should trust frontier labs, because they’re going to capture and then benefit from customer usage data, culminating in them entering competition with those same customers.
2. Everyone is upset frontier labs because they’re overcharging for tokens and the ROI on those tokens isn’t materializing.
3. America needs compelling, powerful open-source models to compete with China.
We think the issue of open-source models is somewhat independent (and in line with what we discussed last week: the real questions to us are who builds a computing model and how they profit from doing so. To the extent there is no profit to be had, or not enough to cover the very high data + training costs that would go into a competitive open-source model, then this is basically a call for government intervention.)
But the other two claims are interesting because, as Claude would probably put it, they exist in tension. Taken together, they’re an assertion that nobody can get enough value from AI but, at the same time, a lab with access to usage data can rip off any business model it lakes. While you can believe this combination, doing so takes a very specific and narrow set of beliefs about both the state of the AI world and also the willingness of closed weight labs to play contractual hardball. Let’s get into why – or why not -- those beliefs might hold.
First, if you haven’t read reports like like this or this, we think they’re worth checking out – the tl/dr is that spend on synthetic data and training is tremendous. Even Anthropic, which likely has the single most profitable model (though we’ll wait for the S1 and GAAP numbers to confirm), is running under 10% margins once you account for training.
With those reports out of the way, the second claim is challenging. A poor value prop on closed source models is either driven from underlying model inadequacies or excessive value capture from the supplier chain. Training costs (and their margin impact) are essential for making models easier to use, but the only layer of the stack with unequivocally strong margins is hardware – Nvidia, memory providers, and the like – who are every bit as upstream of the open weight model usage that Karp is promoting.
That leaves model usability as the primary ROI suspect. We don’t disagree with this, and the number of harness providers entering the market suggest that there are clear challenges. But the foundation labs themselves are not exempt from these challenges. Claude Code and Codex are the two most powerful lab tools, and code capabilities have been on the receiving end of a good portion of the billions of dollars invested in model capabilities. We’d argue that the labs’ advantage here likely came more from non-customer data sources; if one or two large customers were enough to change the direction on model quality, then Google/DeepMind, with its in-house crew of ~80k software engineers, would be seeing different outcomes for Gemini’s coding rank.
We suspect the reason for this is that when you combine the level of data anonymization demanded by enterprise ToS (or even, presumably, Google’s in house lawyers, who do not want to open source their ads business) under which most heavy-weight usage occurs with the usage volumes of actual enterprises, there simply isn’t enough training signal versus what you get from vendors like Mercor. As such, advancements are coming more from direct research and data investments than usage harvesting, though usage signals remain valuable for labs to understand if their changes are directionally correct in real world environments.
To be fair to Karp, there’s a difference between customer data being necessary to enter a space and compete versus customer data being sufficient. Spend on Mercor and Surge very strongly implies that it is far from sufficient. But it may still be necessary, especially when you think about making a model that performs well under a variety of harness or prompting regimes. The irony there is that aggregated usage data is likely valuable primarily for extending usability of the model, rather than enhancing core capabilities.
It’s also possible to believe that usage data is far more valuable, but that belief requires believing that the labs would actively violate enterprise agreements in order to directly capture data. We think this is unlikely; there are too many enterprise agreements and too many lawyers opposed to it. And if people were willing to trust contracting to protect their data from cloud providers, you have to ask why they won’t (eventually) trust it to protect their data from cloud providers hosting LLMs, ala AWS or GCP.
Regardless of how you view the arguments Karp made, we think they’re a sign of things to come. Anxiety around AI’s impact and ROI is growing everywhere from the board room to municipalities evaluating datacenters. Buyers and builders alike will benefit from transparency around benefits, costs, and timelines.



