SOVEREIGNTY, CONTROL, AND THE UNRESOLVED BET ON PREMISES

The early period of enterprise generative AI adoption could be framed as a rush in which companies set aside some of the ordinary rules of good practice, treating speed as a necessity that excused thinking properly about where proprietary data went once it was fed into third-party models, or who actually governed it once it arrived there.

Now it appears that early reading was naïve, and there will be companies that wish their understanding of their relationship with frontier LLM labs was different. What was actually being handed over was a measure of control over product development, customer relationships, and the accumulated process knowledge that has always been the cornerstone of competitive advantage – these have now been transferred to a party whose commercial interests were never guaranteed to stay aligned with the enterprise’s own.

Whether the misalignment is an accidental result or intentional may be hidden in the small print and subject to conjecture, but the leverage has likely transferred, and the front door cannot be easily closed.

To some extent, data sovereignty has hitherto been a relatively narrow discussion around the boardroom table, concerned with where servers sit and in which jurisdiction, and what rules apply. It is growing wider, however, with the European Commission’s proposal of a tech sovereignty plan to reduce EU reliance on American Big Tech, while a number of European cities have already quit the use of Microsoft Office and Google Docs.

While the regulatory framework governing this keeps growing, it is best read as a visible but lagging expression of a concern that looms large. Whilst this is clearly important, a bigger and more consequential version has very little to do with geography; it concerns who owns the compute a company’s workloads run on, who controls the model performing the reasoning, and what it means commercially for a company’s own product. This is especially significant at a time when the frontier labs of Anthropic and OpenAI need to create the harness on which to secure their own futures as they look to the eventuality of public listings.

Investors are constantly evaluating the landscape of AI investment and whether the level of capital expenditure is justified by the return on investment (ROI). Reading the tea leaves of the current climate is difficult. Recently, the hardware trade is proving volatile, which is a direct reflection of the concern over peak investment. We have been here before, but at this juncture there are extraordinary profits on the table waiting to be harvested, and little on the charts that suggest where the support may be.

It is unlikely that the hyperscalers take their collective foot off the pedal, but their rounds of equity and bond issuance to fund expenditure is telling. At the same time, the rate of change in investment will be scrutinised and may challenge the straight-line extrapolation in spending expectations some might imagine, especially when bottlenecks are raising the costs of the build out.

But whilst this presents a clear problem, the potential for a shift in the ‘what is done where’ model is without a doubt an additional headache. For investors, the conundrum surrounding ROI and whether we are already past peak and on the slippery path down is being complicated by the growing yet unresolved question of data sovereignty and on-premises AI using open-source models.

Nvidia sits underneath almost the entire industry as the dominant supplier of the compute on which frontier models rely and whose customers have gone all in with the Nvidia GPUs for training. But alongside this, the company has been quietly developing its increasingly capable family of open weight models under the Nemotron name, moving from smaller, more tentative releases toward considerably more serious, frontier-adjacent versions.

Despite this being a competitive threat and at odds with Nvidia’s customer, the narrative is straightforward: an American open weight alternative is needed because Chinese labs, rather than any American frontier lab, is becoming the default answer to any on-premises or DIY initiative. This may well be the whole explanation, although it could be only part of it, with Nvidia’s largest customers also pursuing custom silicon of their own specifically to reduce their dependence on Nvidia hardware and lower their costs.

Should, as is likely, Nvidia’s response be one to protect its own position in the stack, its open-source models provide increasingly capable ‘Made in the USA’ capability to companies that want to own their AI infrastructure. For Nvidia, more LLM labs is better, and its strategy lies somewhere between serving its core customer and providing insulation in any forthcoming capex slowdown or competitive threat from ASICs. In the medium term, this is especially valuable to Nvidia as it looks to incorporate Groq for inference.

If a shift of this kind does gather pace, it is unlikely to resemble every enterprise suddenly operating its own frontier-scale model. Very few organisations have the capital, the talent, or the genuine need to work at that level, so the more plausible shape, where it emerges at all, is tiered rather than binary.

The hardest and highest-value problems, the ones where an error is genuinely costly, continue to run against a frontier model rented from one of the leading labs. The bulk of routine, high-frequency work, which makes up the majority of what most organisations actually ask their AI systems to do, moves toward a cheaper open-weight model running on infrastructure the company owns outright, often with the frontier model retained only to check or audit the work of the cheaper one operating at volume – and token economics reinforce this.

Increasingly token usage is acting as a barrier to AI deployment, with rented frontier model prices making high usage unsustainable and which may be seen to hold a company’s own AI aspirations in check. The economics change dramatically when using owned hardware and running an open weight model where token cost per additional query collapses toward zero once the initial purchase is made. Companies are already walking away from the frontier when faced with consumption-based pricing.

Turning the boring and high-volume tier of a company’s AI workload into a rational economic decision rather than a purely defensive one about control may seem obvious, but it is not yet central. There may be other reasons for this given concerns surrounding where the bots are, where they have been, and what they have changed, but this concern will give way as companies begin to trust.

The alignment of sovereignty, control and unit economics would make a shift of this kind durable rather than merely fashionable, but it is also possible that frontier model economics improve faster than the on-premise case requires, or that contractual guarantees, private inference arrangements, or dedicated capacity from the labs themselves close the gap enterprises currently feel between their own data and a third-party model, making ownership unnecessary.

None of this necessarily implies that hyperscale capital spending stops, only that it is worth treating as a considerably less uniform story than current pricing across the sector may suggest.

The announcement by Meta that it was looking to sell compute was initially seen as an admission that it had over-invested and that its internal demand fell short of the capacity in place. This perception caused havoc in technology hardware companies and reinforced the increasing negativity toward the group.

However, what was conveniently forgotten in the stampede for the door was Meta’s prior announcements of compute contracts with others. To recap: in March, Meta signed a US$27 billion cloud service deal with Nebius (the first US$12 billion firm), then in April Meta signed a US$21 billion expansion of its agreement with Coreweave, and just recently it signed up with Crusoe, a smaller Neo, for 1.6GW of capacity. Those are not the actions of a company with too much compute – far from it.

Recently it was reported that Google has been forced to cap Meta’s use of Gemini compute given Meta wanted more than Google could supply. This has forced the company to limit its token use which, for a company spending upward of US$145 billion in capex over 2026, is astonishing. Whilst it may appear an oversimplification, the more obvious conclusion is likely the pervasive one – that Meta has been working to monetise some of its older capacity and generate an ROI on it.

While there is the potential that Meta has built too much and consequently is selling excess capacity, it seems far more likely that it has been deliberately and diligently triaging its compute, working to clear older-generation capacity to make room for the hardware it wants for its own frontier ambitions.

Read one way, Meta’s monetisation of its older compute looks like ordinary capital recycling in the middle of an arms race, the kind of housekeeping any well-run business would undertake. Read another way, it may be an early signal that not every participant currently building at this scale will end up needing everything they have built, and that capacity assumed to be permanently scarce has a way of becoming available once a competitor’s own ambitions fail to pan out.

That would carry real consequences, but currently it may also seem a long bow to draw given we are in the early stages of adoption, although a shift of demand toward more distributed desktop and edge-scale hardware would also carry rather different implications for the chip and component supply chain than a build out concentrated in a small number of hyperscale clusters.

A harder question sits underneath, and that is the potential valuation attached to both Anthropic and OpenAI given they currently assume a world in which enterprises continue routing an expanding share of their workload through their frontier models. Does this erode their ability to raise capital, or keep their inexorable march to the front in check, and see one of them falter?

For the time being, whether the need for data sovereignty ultimately drives owned infrastructure to overtake that of the clouds in AI compute is uncertain, if not somewhat unlikely, but it is certainly being viewed as a vulnerability at a time of great nervousness. However, there are clouds on the horizon, and they should not be ignored even if a definitive answer is difficult to distil.

Alongside the rising costs of the AI build out, itself driven by the rise in bottleneck economics, the open model and on-premises debate may weigh on hardware company share prices for a while yet, unless the next round of earnings releases can calm frayed nerves. But the fundamental question about who ends up controlling the means of production for a company’s own intelligence is a significant debate, and one where Nvidia has not fully revealed its hand. Remembering the dance may serve well, but perhaps it is also better to continue the dance nearer the door?

Tim Chesterfield is CIO of the Perpetual Guardian Group and the founding CIO and Director of its investment management business, PG Investments. With $2.8 billion in funds under management and $8 billion in total assets under management, Perpetual Guardian Group is a leading financial services provider to New Zealanders.

Disclaimer
Information provided in this publication is not personalised and does not take into account the particular financial situation, needs or goals of any person. Professional investment advice should be taken before making an investment. The information provided in this article is not a recommendation to buy, sell, or hold any of the companies mentioned. PG Investments is not responsible for, and expressly disclaims all liability for, damages of any kind arising out of use, reference to, or reliance on any information contained within this article, and no guarantee is given that the information provided in this article is correct, complete, and up to date.

You might also like:

Top