After the Interview: What I Wish I'd Had More Time to Say About the AI Grid
The Digitalisation World interview covered the AI Grid concept in eight minutes. Here's what that format couldn't fit: the three questions I get asked most often afterwards, and why the answers matter more than the headline.
The Digitalisation World interview ran to about eight minutes. That was enough to land the core idea — inference isn’t training, CDN architecture applies to AI, distributed GPU compute changes the economics — but not enough to deal with the objections that come up whenever I discuss it in person.
The same three questions recur afterwards. They are more useful than the headline because each exposes a different constraint on where inference should run.
“Isn’t this just marketing for edge vendors?”
It’s the right challenge. The AI Grid framing originated with NVIDIA’s reference design, and every CDN and edge provider with a GPU roadmap has since adopted the language. I would be sceptical of vendor-led claims about when and why the grid matters.
I would start with latency, because it is real: real-time video personalisation, gaming AI that has to land inside a frame budget, customer-facing agents where a 400ms response in Tokyo versus a 40ms response in Frankfurt is the difference between a product that feels instant and one that feels broken. For those workloads, centralised inference isn’t just slower — it’s structurally more expensive, because you end up over-provisioning frontier capacity to compensate for a problem that a small model at the edge solves outright.
Latency is not the whole test. The consolidation economics apply even to workloads with no real-time requirement. One GPU card, one loaded model, one warm KV cache, serving many concurrent users is materially cheaper per token than many cards each carrying their own model instance and cold context. Fragment inference across teams, regions, or consumer instances — each spinning up its own deployment — and you multiply the fixed cost of model loading and context initialisation across every instance. Consolidation onto shared, well-utilised infrastructure recovers that waste. The grid is the mechanism that lets you consolidate without sacrificing the geographic or jurisdictional constraints that fragmentation was originally solving.
The test I apply is not latency alone. It is whether the current deployment pattern is paying the fragmentation tax — in GPU memory, in cold-start overhead, in underutilised capacity — when a better-placed, shared deployment would serve the same workloads more cheaply. That applies to internal copilots and batch pipelines too.
“What does ‘distributed footprint’ actually mean in practice?”
When I said in the interview that the commercial advantage tilts towards whoever already owns a distributed footprint, I meant a specific claim about capital structure — not a vague observation about scale.
Building a global network of edge locations from scratch takes a decade and tens of billions in capital expenditure. The fibre, the facilities, the power contracts, the peering relationships — none of that is fast to acquire. CDN providers and edge networks spent the last twenty-five years building exactly this infrastructure to serve web content and video. The AI Grid upgrade cycle, for them, is adding GPU compute to locations that already exist and are already connected.
For a hyperscaler starting from a handful of large regions, the path to the same geographic coverage is genuinely long. You can’t buy your way to thousands of edge PoPs quickly — the physical infrastructure constraints are real.
For enterprise buyers, I would separate GPU capacity from placement. The relevant question is which provider can serve users in the regions that matter to the product, at the latency the product requires, without routing every request back to a central cluster. Those are different questions, and they don’t always have the same answer.
“How does this interact with data sovereignty?”
This is the question I get most often from UK and European enterprises, and it is the one the eight-minute format could not do justice to.
Sovereignty is a third dimension in the routing decision — alongside latency and model capability. An orchestrator routing inference requests needs to know which node is closest, which node can run the model, and which nodes are jurisdictionally eligible for the data.
For many enterprise workloads, the sovereignty constraint is the binding one — though it’s worth being precise about why, because the law and the policy are two different things and people routinely conflate them.
UK GDPR does not draw a hard border at the UK and EEA. A restricted transfer can be perfectly lawful: to a country covered by UK adequacy regulations, or on the back of an appropriate safeguard such as an IDTA, the Addendum to the EU SCCs, or binding corporate rules. What it does is attach conditions, paperwork, and a transfer risk assessment to every route that leaves. Plenty of organisations then go further than the law requires and adopt a flat in-jurisdiction policy for customer data — sometimes because the assessment is more expensive than the alternative, sometimes because a regulator, a board, or a large customer’s contract asked for it. A financial services firm or an NHS-facing healthcare provider is more likely than most to land there.
Either way, the routing constraint is real. It usually originates in an organisation’s own data policy rather than in a categorical legal prohibition — and that matters, because a policy is something an organisation can revisit with its DPO, while a border is not. The grid does not remove the constraint, whichever it is. It puts more nodes inside the eligible set, so compliance and performance do not have to be traded against one another.
The practical implication is that, when evaluating AI Grid deployments, the density of edge nodes within the jurisdiction matters as much as the total global footprint. A provider with two UK nodes and a thousand US nodes is less useful to a UK-regulated enterprise than one with twenty UK nodes and two hundred globally. That is a different evaluation criterion from the one most infrastructure conversations default to.
Pole position is an option, not a deployment date
The interview framed the AI Grid as an infrastructure question. I think the more useful version is strategic.
I use the CDN analogy because we know how that story ended. When CDNs became the default delivery mechanism for web content and video, the competitive advantage shifted from “who has the biggest origin server” to “who has the most distributed delivery network.” The companies that had already built global edge infrastructure captured that shift. The ones that tried to build it after the fact mostly didn’t.
AI inference is following the same trajectory. Distributed inference will matter because the physics, the economics, and the user expectations are the same. For enterprises, the decision is whether to position AI infrastructure to take advantage of that shift or retrofit it after the fact.
That is the pole position the interview title was pointing at. It is not about being first to deploy a grid. It is about making infrastructure decisions now that do not foreclose the option later.
The full interview is on Digitalisation (opens in a new tab). The AI Grid strategy page covers the three-tier architecture in more depth, and the routing calculator lets you model the latency, sovereignty, and cost constraints for your own workloads.
Source: https://bradshaw.cloud/writing/ai-grid-pole-position-followup/