Markets, strategy, implications

Tier-1 Source Links

Signals

The week's tier-1 sources, newest first

The Signals feed surfaces the primary sources behind each edition — the tier-1 posts the pipeline ingests, linked straight through.

AgentPulse · living reference

The agent economy, live

What has to exist for AI agents to actually work, what is missing right now, and why almost everyone is watching the wrong part of the chain.

Live Updated 16 Sep 2026 · thesis live since 13 Sep 2026 · corrected 3 times

Of the five physical things the agent economy runs on, four are tied so closely that none can be called the bottleneck. Only one of them has no way out — and this August it hit a ceiling the industry does not expect to lift before the end of the decade. Everything below explains why that matters, starting from the beginning.

This page holds two kinds of content and keeps them apart on purpose. What explains how things work barely changes. What says where we are changes every few weeks, is marked in violet, carries its own date, and opens up to show what it said before and why it changed.

You need to know nothing about the subject. Every figure has a source at the end.


Part I · The idea

What just became possible

Software has always done what you told it. You specified the steps; it ran them. If you left out a step, it failed.

That changed. A program can now be given a goal instead of an instruction, and work out the steps itself. "Find me a flight on Thursday that gets in before the meeting" is not a set of instructions — it is a description of what success looks like. The program decides the rest: which airlines to check, what counts as too tight a connection, whether to warn you that the cheap one lands at a different airport.

Call that program an agent. What makes it an agent is not that it is clever. It is that it makes decisions you did not specify.

And that changes what a person can do with it. You can hand it a task and stop watching. That sounds small. It is the whole thing — because the moment you stop watching, you are no longer using a tool. You are delegating.

And delegation is where it gets interesting

Here is the part that turns a feature into an economy: a thing that can be given a task can also give one.

An agent booking your travel can hire another agent to check visa rules. That one can pay a third for access to a database. None of those transactions has a person in it. When two agents strike a deal without a human in the loop, everything that comes with commerce arrives at once — prices, contracts, counterparties, deadlines, disputes, and fraud.

That is the agent economy. Not AI that is smarter. AI that transacts.

Why it is harder than it sounds

We already know how to delegate. We do it constantly — to employees, contractors, lawyers, builders. It works, and it works because of five things we never think about, all of which assume the person we are delegating to is a person.

They can be punished. Sued, fired, fined, struck off, jailed. The possibility of consequence is what makes someone think twice before cutting a corner.

There is only one of them. Identity is scarce. Reputation only means anything because you cannot escape yours.

They are slow. A person can only do so much damage in an afternoon. Every limit we write — spending caps, approval thresholds, notice periods — is calibrated to human speed without ever saying so.

They can be asked afterwards. What did you do, and why? The answer is evidence. It can be tested, and it can be held against them.

They know when they are out of their depth. Imperfectly, but usually.

Now take each one and point it at an agent.

It cannot be punished. You can delete it. It does not mind. There is no version of consequence that reaches it, so there is no version of deterrence.

It can be copied. Perfectly, a million times, for nothing. Attach a reputation to an agent and it is worthless: burn one, start the next. Scarcity of identity was doing invisible work, and it is gone.

It is not slow. Thousands of actions a second. A spending limit with no rate limit is not a limit.

Its account of itself is not evidence. Ask an agent why it did something and you get a plausible explanation, generated now, of a decision made earlier. And because these systems are not deterministic, "run it again and see" does not settle anything either.

It does not reliably know when it is out of its depth. This is the one that catches people out, because a confident wrong answer looks exactly like a confident right one.

So what can actually be delegated?

The honest answer today: whatever is cheap to be wrong about.

Drafting, summarizing, searching, sorting, monitoring, first passes at almost anything. All tasks where a mistake costs a little time and nothing else.

What cannot be delegated is not what agents are bad at. It is what needs one of those five guarantees — anything where being wrong costs money you cannot get back, commitments you cannot undo, or harm to someone who then has nobody to take it up with.

That is worth saying plainly, because it is usually said wrong:

The limit on delegation is not capability. It is recourse. Agents can already do plenty of things we do not let them do, because there is no way to be made whole when they get them wrong.

Which means the work of building the agent economy is not mainly about better models. It is about building substitutes for five things a person gives you for free.


Part II · What has to exist

Those substitutes do not sit in one place. They stack.

At the bottom is the physical world — the chips, the memory, the buildings, the electricity. Above it sits a layer of institutions that mostly does not exist yet: identity, memory, payments, autonomy, governance. Each of those has a problem that is genuinely unsolved, and each one, when it is solved, will send a bill downstairs.

LAYER 0 · WHY DEMAND GROWS Demand for intelligence HOW MANY AGENTS × HOW MUCH EACH DOES × HOW MUCH IT THINKS × HOW MUCH IT CHECKS ITSELF LAYERS 2 · 3 — THE INSTITUTIONS THAT ARE MISSING Identityand trust Memoryand context Payments andsettlement NO PHYSICAL BILL Autonomyand control Governanceand accounts EVERY DECISION UP HERE HAS A PHYSICAL BILL DOWN THERE LAYER 1 — THE PHYSICAL WORLD · LOWER = BINDS SOONER Chipassembly 1.57 TIGHT Rentingspace 1.62 TIGHT Memory(HBM) 1.70 TIGHT Poweredcapacity 1.80 TIGHT Siliconwafers 2.73 SLACK NOT A CONSTRAINT — IT RELIEVES, IT DOES NOT LIMIT The software that serves the models · every gain here subtracts demand from everything above

Click any block for its detail · the numbers are explained in Part III

Identity — the copy problem

Proving who you are was solved decades ago. That is not the problem.

The problem is that an agent never acts for itself. It acts for someone. So the question is never "which program is this" — it is whose authority is it acting under, and how much of it.

And the copy problem makes the obvious answer useless. If identity is a property of the program, it can be duplicated a million times over and the word stops meaning anything. Whatever identity is for an agent, it cannot be a name. It has to be a grant: this may spend up to this much, on this, until this date — in a form a stranger's system can check without phoning anyone.

Then comes the hard half, and it is stranger than it looks.

Our entire apparatus for withdrawing authority runs on threat. Break the terms and something bad happens to you, so you do not. An agent cannot be threatened. So authority cannot be withdrawn by punishing — it has to run out on its own.

Permissions that expire, rather than permissions that get revoked. That is a genuinely different way to build a system, and almost nothing is built that way today.

Its physical bill: silicon that can prove which code is actually running — attestation. Specialised compute, not generic compute.

Memory — what it keeps, and who pays for it

An agent that forgets everything between one task and the next is useless. One that remembers everything is expensive and dangerous.

The institutional questions are what persists, who can read it, and who answers when it leaks. Those are real and mostly unanswered.

But this is also the block where the whole structure of this page first shows itself, so it is worth slowing down. Remembering means holding data in fast memory, and fast memory is the scarcest thing in the entire chain.

Which makes an apparently abstract design decision — how much context should this agent carry? — into something very concrete. It is a hardware purchasing decision. It always was. Nobody writes it down that way.

Its physical bill: HBM. The same tight block that appears further down this page.

Payments — a thousandth of a cent, a thousand times a second

An agent that cannot pay for anything cannot do much. It can read, but it cannot hire, buy data, rent compute, or settle with another agent.

The usual explanation for why this is hard is that banks are slow, and it is wrong. Instant payment systems exist and work — FedNow, SEPA Instant, UPI. Speed is not the obstacle.

Two other things are.

The amount. Every rail in use has a floor: interchange, minimum fees, the cost of clearing a single transaction. Below some amount, processing a payment costs more than the payment. Agents want to pay per API call, per thousand tokens, per second of compute — amounts far under any floor that exists. Not slightly under. Orders of magnitude under.

The authorization. Every rail assumes an account holder who is a legal person and who authorizes with intent. An agent making ten thousand decisions a second cannot be that. "The human approved a budget beforehand" is not the same thing as authorization, either legally or operationally, and the difference shows up precisely when something goes wrong.

So what is needed is three things together: amounts far below a cent, settlement that is final immediately, and authorization that is programmatic and bounded rather than personal and open-ended.

That combination is why crypto rails keep coming up in this conversation. Not because they are better in general — they are not — but because stablecoin rails and the protocols being built on them are the only place where all three are being attempted at once. It is worth being clear-eyed about how finished that is: the leading payment protocol's claim of final settlement is conditional on institutions in the critical path, and the leading authorization standard only added machine-initiated payments recently. Half-built, moving, and the only thing moving.

Its physical bill: almost none. A payment is a ledger entry, and ledger capacity is abundant. This is the only block on the map with no meaningful physical cost — and that absence is a finding, not a gap in the research.

Autonomy — how much, and how fast

Two different things get talked about as one.

How much an agent may commit is a question about size. How quickly it may commit is a question about rate. Every control we know how to write handles the first and ignores the second, because for people the second never mattered.

A thousand-euro limit means nothing if it can be spent a thousand times a second.

That is not a detail to be patched later. A size limit without a rate limit is not a limit at all, and almost every permission system in existence is a size limit.

Its physical bill: reasoning compute. The more autonomy an agent has, the longer it deliberates before acting — and deliberation is the fastest-growing line in the whole consumption bill.

Governance — who answers when it goes wrong

Because the agent cannot be punished, responsibility climbs to whoever deployed it. That sounds like it settles the question. It opens four.

The chain has five links. An agent runs on one company's model, inside another company's framework, on a third's infrastructure, acting for a user, calling a fifth company's tool. When it causes harm, there is no settled doctrine for dividing the blame. Ordinary software has contracts and product liability; here the chain is longer and no single link controls the behaviour that emerges.

Proving what happened is unusually hard. Months later, to someone who was not there, you need a record that exists, that could not have been altered, and that can be read. And the normal fallback is unavailable: because these systems are not deterministic, "reproduce it and show me" proves nothing. The record is the only evidence there will ever be, which means it has to be kept always, not from the moment you start to suspect.

Undoing things assumes human timescales. Chargebacks, cooling-off periods, statutory notice — all of it exists because a person takes days to notice a problem. An agent can enter and settle thousands of commitments before anyone looks.

And the genuinely new one: behaviour changes without anyone changing it. The model is updated. A tool it calls starts answering differently. The context is not what it was. "It was compliant when we deployed it" quietly stops being true while nobody touches anything.

That last one has a consequence that runs through the rest of this page. It means governance cannot be a certificate issued once. It has to be continuous — every claim already made has to keep being checked against what becomes known later.

Its physical bill: surveillance compute, and it is the least obvious and largest on the map. Watching what you already claimed costs more than claiming it, and it grows with the number of live commitments you are carrying, not with what you produce. It is a stock, not a flow.

Two more, named and left out

There are two further blocks that belong in any honest map of this and are not measured here: negotiation and coordination — how agents reach agreement without a person arbitrating — and disposition — what an agent's standing tendencies are, and whether they are stable.

They are left out for one reason, and it is the same rule applied everywhere else on this page: their physical cost has not been characterised, so putting them on a map that is about physical cost would be decoration. They come back when there is something real to say.


Part III · The loop

Everything above is paid for downstairs

Read back through the five blocks and look only at the last line of each.

Identity needs silicon that can attest. Memory needs the fastest memory there is. Autonomy needs reasoning compute, and more of it the more autonomy you grant. Governance needs surveillance compute that accumulates rather than flows. Payments need almost nothing, which is itself worth knowing.

So the institutional layer is not floating above the physical one. It rests on it, and it presses down.

And here is the turn, which is the reason this page exists:

Solving the institutional problems does not relieve the physical pressure. It creates it. Every identity problem that gets fixed means more agents that can act. Every agent that acts consumes memory, compute and power. The better the institutions get, the harder the floor gets pressed.

There is exactly one thing pushing the other way: efficiency. Models get cheaper to run, attention gets cheaper to compute, the same work costs less every year. If efficiency outruns adoption, none of this binds and the whole question is academic.

It has not been outrunning it. Which is why the rest of this page stops talking about institutions and starts measuring the floor — how much room is left in each physical thing, how fast each one can be expanded, and which of them runs out first.

That is a question with a number, and the number has moved twice this month.

What is downstairs

Five things carry that bill. They get talked about as if they were one — "chips", or "compute", or "the data centres" — and that is the single biggest source of confusion in the public argument about AI.

They are not one thing. They are five separate industries with five different clocks, and it is the difference between those clocks that decides which one runs out first.

Chip assembly

A modern AI chip is not a chip. It is several pieces mounted together: the processor on one side, stacks of memory beside it, all sitting on a shared slab of silicon that lets them talk to each other at enormous speed. Putting those pieces together is its own manufacturing step, in its own factories, and it is the narrowest point in the whole chain. Almost all of the world's capacity for it goes to AI.

Expanding it takes four things and each one takes time: cleanroom space, the bonding machines themselves — which have their own waiting list running into quarters — installing them, and qualifying them.

Then there is yield, which is where it usually goes wrong. The pieces heat up and expand, but silicon and substrate expand by different amounts, so the assembly warps. A warped package is scrap. That problem, not any shortage of chips, is what delayed a whole generation of NVIDIA accelerators in 2024.

Time to expand: 12 to 24 months.

Memory

The memory that feeds an AI chip is stacked — several layers of memory one on top of another, connected by holes drilled straight down through the silicon, and placed right next to the processor that uses it. That closeness is the point. It is what gives the bandwidth.

Expanding it needs two different capacities at the same time, and they are not the same business. You need memory fab capacity, which is a two-to-three year investment cycle. And you need stacking and qualification capacity, which is a separate thing entirely — and whose yield gets worse the higher you stack. Twelve floors is much harder than eight.

There is a detail almost nobody mentions: a wafer given over to this kind of memory produces far fewer sellable bits than one making ordinary memory, because the stacking eats area. So growing the supply of AI memory reduces the supply of the normal kind.

Time to expand: 18 to 36 months — the longest of the tight group.

Renting space

This one is not physical at all, which is why it confuses people. The buildings exist and they have power. They already have an owner.

Whoever expects to receive chips a year from now signs the lease today, because waiting means there will be nothing left. So space gets committed long before it gets filled — nine to nineteen months between signature and actual load, with the result that a market can be fully leased and half empty at the same time. One company had barely a fifth of what it had contracted actually running.

What matters about this one is that nothing has to be built for it to ease. You wait for contracts to turn over. It is the only one of the five whose relief is measured in months rather than years.

Time to expand: 9 to 19 months — a contract cycle.

Powered capacity

This is the one the public argument is loudest about, and it is four problems wearing one name. Only one of them is the building.

You need land near enough electricity. You need a place in the queue to connect to the grid, which in some regions runs past three years. You need transformers, whose delivery times have reached five years. And then you need to build the thing and cool it.

The grid queue is the part that governs, and it varies brutally from region to region — some areas are effectively closed while others have room. Any national average hides that completely.

One thing worth holding onto when you read that data centres sit half empty: a serious one is built with redundancy, backup systems that exist precisely so they are never used. Part of that idle capacity is design, not waste.

Time to expand: 2 to 7 years — the slowest thing on the map.

Silicon wafers

The famous one. Building a leading-edge chip factory costs tens of billions and takes over three years, which makes it sound like the obvious bottleneck.

It is not, and the reason is simple: AI is a small customer. It consumes something like an eighth of the world's most advanced wafer capacity. Phones and everything else take the rest.

So nothing has to be built. Capacity has to be reallocated, by outbidding the phone maker — which is much faster than putting up a factory, though not instant, because changing what a fab makes means requalifying processes and the commitments are signed quarters ahead.

The thing to watch: if AI ever took 80 to 85% of a leading node, that reallocatable slack would be gone and wafers would start to bind like the rest.

Time to expand: 24 to 36 months for new capacity — but months to reallocate what already exists.

Five industries, and the fastest of them eases in nine months while the slowest takes seven years. That spread is not a detail. It is the reason one of them, and not the others, is about to be the thing that runs out — and working out which one is a question with an actual answer.


How to read it

Five ideas and that is all

This section does not change. It is the mechanism, and the mechanism is the same today as it will be in a year.

1 · What matters is what is left, not what there is

What matters is not how much electricity a country produces, but how much it produces divided by how much it uses. A two-hundred-table restaurant with every table taken is exactly as full as a ten-table one with ten taken. That ratio is what we call slack.

2 · There are two clocks

Demand grows at one rate and capacity at its own. Only the difference matters. If capacity is slower, the margin shrinks to zero. If it is faster, that thing stops being a problem forever.

3 · The winner is whichever hits zero first, and it is not the slowest to fix

Building a power plant takes years; expanding an assembly line takes months. Intuition says electricity must therefore be the problem. But how long something takes to fix tells you nothing about when it breaks. Slow with plenty of margin holds for years; fast with no margin breaks tomorrow.

4 · How much demand grows barely matters

Demand rises for everything at once, so it shrinks every margin by the same proportion — and that does not change which one is smallest. Demand is the tide; slack is the seabed. The tide decides when the boat runs aground, the seabed decides where.

That holds only while every unit of demand arrives with the same recipe: so much chip, so much memory, so much power, in fixed proportion. It has stopped doing that — which is what the section after next is about.

5 · Build time is the lag on the response

When something runs short its price rises, and price attracts money to build more. But that money takes exactly as long to become capacity as the construction takes. What expands in three months cures itself; what takes four years cannot respond within four years.

The bottleneck moves toward whichever link has the longest build time, and gets stuck there.


What the data says

Four tied and one alone in the clear

Oldest data it rests on Dec 2025
Live Reviewed 13 Sep 2026 · changed once

Four constraints land between 1.57 and 1.80 on the tightness index. The fifth — wafers — sits at 2.73, 52% above them[6][7].

Within the four there is no defensible order. Powered capacity is the extreme case: its range runs from 1.49 to 1.99, which overlaps all three of the others entirely. It could be the tightest or the slackest of the four, and the available data cannot tell you which.

Naming one of them "the" bottleneck would be inventing a precision the data does not have.

Why sites look full while sitting half empty is documented: companies sign the lease long before they have the chips, and the space takes between nine and nineteen months to fill from signature[3]. One company had only 21% of what it had contracted actually running[4]. The head of one of the largest technology companies in the world described the problem not as a shortage of chips but as a shortage of warm shells to plug them into[2].

And the bottleneck has moved before

The most solid thing on this map is not a prediction: it is an observation somebody else already published.

"As TSMC expanded its assembly capacity through 2025, this bottleneck eased… As assembly loosened, memory tightened."

Epoch AI, March 2026 [6]

That is the fifth idea happening in the real world: one constraint eased, the one next to it absorbed the tension, and the pinch point moved. There is a whole sequence documented since 2023, hopping from link to link.

Prediction Issued 13 Sep 2026 · resolves Mar 2028

Over the next eighteen months, of the five constraints only assembly could begin to ease — its expansion time is 12 to 24 months, and only the bottom half of that range fits inside the window, and only if the build was already under way. The other four do not make it. The tight group stays tight and the bottleneck rotates inside it.


The recipe is changing

The last thing we measured, and why we are not using it yet

Live Measured 14 Sep 2026 · new

There is a way to see this with no arithmetic at all. NVIDIA sells two consecutive generations of the same system, the B200 and the B300. They do exactly the same amount of computation: 2,250 trillion operations per second, the identical figure on both. Memory goes from 192 to 288 gigabytes[10].

Fifty per cent more memory for the same compute, in a single generation. Nobody asked for more computing power: they asked for more room to keep things in.

And it is not one company's call. NVIDIA's GB300 and AMD's MI355X ended up on the same three numbers — 288 gigabytes, the same compute, the same 1,400 watts[11]. Two design teams that do not talk to each other landed on the same point, which suggests the ratio is not something a vendor picks: it is set by what the work needs.

And the work that needs it is agents. Almost everything that makes an agent different from a chatbot is memory-hungry and compute-cheap: keeping a long conversation alive, taking fifty steps instead of one, remembering what it did last week. A chatbot answers and forgets. An agent has to keep everything it has already done present.

What this would change, and why we have not done it

Feed this measurement into the model and memory stops being tied and becomes the tightest constraint, alone. It would be a far cleaner headline than the one we have.

We have not done it, and the reason is a problem with our own numbers. Of the five figures we compare, memory's comes from a bank's forecast that already includes an estimate of future demand; the other four are supply figures to which we add the demand estimate ourselves. Those are two different things sitting in the same table. Applying the correction to four and not the fifth — using as the criterion the very problem we have not solved — would produce a confident and probably wrong answer.

So the measurement is done, published here, and waiting. When all five figures are built the same way it gets applied, and this page will say so.


Ways out

Four are tied. Only one has nowhere to go.

Everything above measures how tight each link is today and how fast its supply is growing. That is a photograph, and photographs move — this one has moved twice in a week.

There is a different question that does not depend on any of our numbers: when a link runs short, how many separate ways are there out? Not how long the fix takes — how many independent fixes exist at all. Answer that and you learn something the index cannot tell you, because it survives being wrong about the index.

The fiveHow each link gets relieved, and how many routes it has

Renting space — Nothing is built and nobody is asked. Leases expire, hoarded capacity comes back to market, and the squeeze eases on its own within a contract cycle. Its relief needs only time.

Silicon wafers — AI uses 11–12% of the most advanced capacity; phones use most of the rest. Relief is outbidding the phone maker, not building a fab. One route, and it is a price.

Chip assembly — Several technology variants that are not the same product, more than one company able to do it, and something the others cannot claim: it has already been relieved once, in 2025, and a third party documented it happening[6].

Powered capacity — The slowest clock on the map, two to seven years. But it is slow in a thousand places at once: generate on site, build in a region with a shorter grid queue, shift load, go abroad. Many routes, all slow, all independent of each other.

Memory — One route. Stack the chips higher. That is the whole list.

Live Measured 15 Sep 2026 · new

And this August, that one route hit a ceiling with a number on it.

You cannot widen a memory stack: there is only so much room beside the processor, and the industry has been stuck at eight stacks per accelerator for two generations. So the only way to add memory is upward. Upward now has a limit too, and it is 775 microns — about eight times the thickness of a human hair.

The reason is almost comically physical. Under the cooling plate, the memory and the processor are both ground down until bare silicon shows. A processor wafer is 775 microns thick. If the memory stack were taller, it would stand proud of the chip beside it and the assembly would not close. An SK hynix vice-president put it plainly at a conference on 23 August: "that's the kind of limit that we can go up so far, because the logic wafer thickness is also 775 microns"[12].

There is a technique that would buy room back — bonding the layers directly instead of joining them through tiny solder balls, which reclaims the space the balls take up. It is not coming in time. SK hynix has said it will not be ready for the next memory generation, and expects it one generation after that; outside analysts put volume production around 2029–2030[12]. NVIDIA's entire next accelerator generation will be stacked the old way.

Two caps, then. Capped sideways by the space beside the processor, capped upward by the thickness of a wafer, and the tool that lifts the second cap arrives at the end of the decade.

The precise version, because the loose one is wrong

That ceiling does not cap how much memory gets manufactured. Build more plants and more memory exists — that part is only money and time.

What it caps is how many gigabytes fit on one accelerator. Which is exactly the number that has been climbing: the same chip design now carries 50% more memory than its predecessor for identical compute, and that figure has risen every generation since 2022.

So the cap sits on precisely the axis demand is pushing against. That is narrower than "memory is short", and it is the reason the shortage is hard to design around rather than merely expensive.

Which gives the asymmetry, and it is the most useful thing on this page:

Every link in the chain has more than one way out except one. Assembly has already taken its. Rented space frees itself. Wafers need a higher bid, not a new factory. Power is slow but can be built in a hundred places at once. Memory has a single route — stack higher — and this August that route ran into a physical ceiling nobody expects to lift before the end of the decade. It has three suppliers, and all three meet the same ceiling.

Memory is also the only one of the five where every arrow points the same way. Supply is capped in both directions it could grow. Demand per unit of work is the only one of the five that is rising rather than falling. Everywhere else, at least one arrow points toward relief.

Live Reviewed 15 Sep 2026

Two things would break this, and both are live.

A fourth supplier. China's CXMT began small-batch production of stacked memory in August and listed in Shanghai with plans to give a fifth of its capacity to the product line this year[13]. It is building a generation or two behind what a current accelerator needs, and its own schedule has already slipped[13].

But be precise about what a fourth name would change. It would add capacity — more memory made, by more hands. It would not add height. A newcomer meets the same 775 microns as the incumbents and needs the same technique that is not ready. Concentration is the weaker half of this argument; the ceiling is the half a competitor cannot break.

A change in how models are built. This is already happening and it cuts the other way: newer attention designs have reduced the memory a model needs per unit of conversation by a factor of between eight and fifty. Set against that, the move toward models that keep far more parameters resident pushes memory demand up. Right now the two roughly cancel. That is a tie between two live trends, not a law — if either one stops, the answer moves.

One last thing about what this claim is and is not. It rests on no measurement of ours — not the index, not the tightness ranking, none of the figures this page flags as uncertain. It rests on counting routes, and on one number a manufacturer stated in public. It would survive every figure above being wrong. That is unusual here, and it is why it sits near the end rather than the beginning: it is the part of the page most likely to still be true in a year.


Inside each link

The plain-language version of each physical block now lives in Part III, under “What is downstairs”. What stays here is the technical depth — the precise figures, the terms of art, and each block’s clock — for the reader who asks for it afterwards.

The physical world

AssemblyWhat is advanced packaging and why does expanding it take two years?

The shared silicon base the pieces sit on is an interposer, and the assembly step has a name: advanced packaging. It is what separates an AI accelerator from an ordinary processor.

The generation the warping problem delayed in 2024 was NVIDIA’s Blackwell.

Hence the lead time12 to 24 months, and only the bottom half of that range fits inside this prediction's horizon.
HBM memoryWhy is making more memory not enough?

The stacked memory has a name: HBM, high-bandwidth memory — the same term the glossary opens on, with the full technical ladder inside.

Hence the lead time18 to 36 months — the longest in the tight group, which is why the bottleneck tends to settle here.
Renting spaceWhy is everything leased if it is half empty?

This constraint is not physical: it is a market one. The space exists and it has power, but it already has an owner — vacancy in the primary US markets is 1.4%, with under 1,500 MW available[1].

Hence the lead time9 to 19 months, the contract cycle.
Powered capacityWhat does it take to build new capacity?

One nuance worth holding when you read the “42% unused” figure: a serious data centre is designed with redundancy — backup systems that exist precisely so they are not used. Part of that idle capacity is design, not waste.

Hence the lead time2 to 7 years. The slowest thing on the map.
WafersWhy are the chip fabs not the bottleneck?

The precise share behind “something like an eighth”: AI consumes around 11–12% of the most advanced fab capacity.

Hence the lead time24 to 36 months for new capacity, but months to reallocate what exists.

Honesty

What we do not know

Live Reviewed 13 Sep 2026

The number that decides the thesis does not exist publicly. It is the growth of memory supply in 2026, and it only sits in paid reports.

Powered capacity depends on which facility you look at. In AI-dedicated sites, built with very little redundancy because a training run restarts and a bank transfer does not, the margin lands just above the tight group. In conventional colocation with double redundancy it lands clearly below[5][9]. They are two different worlds and they fall on opposite sides.

A correction published the same day. An earlier version of this text said the academic study this figure comes from contradicted itself between draft and published version. That was false: only one version exists, and the two numbers are different scenarios within it, with a stated reason for choosing one as central. The error was ours in reading it.

Five figures that are not built the same way. Memory's comes from a bank's forecast that already nets out future demand; the other four are supply figures to which we add demand ourselves. Comparing them as if they were the same thing is the most serious flaw this map currently has, and it comes before every other one: until it is fixed, no new correction can be applied without risking counting the same effect twice. It is declared here before it is solved, on purpose.

Almost everything comes from interested parties. Fab capacities are declared by the manufacturers[8]; deficits are estimated by banks that trade in those markets[7]. "Sold out" is a commercial claim, not a measurement.

Scopes are mixed. Two figures are United States and three are global. They are compared because there is no alternative, but the comparison is directional.

It is worth being precise about what survives this and what does not. The mechanism — the pinch migrates toward whatever takes longest to expand and stays there — depends on no measurement of ours: it rests on a handover that already happened and that a third party documented. The snapshot does depend on us, and it has already moved once in a single day. The order within the tight group does not hold, which is why it is not claimed. The only thing claimed with numbers is that wafers sit outside the group, and that margin is larger than the error we made.


Sources

Where each figure comes from

Every figure on the map carries its measurement date and its strength: measured (counted by a primary source), derived (calculated from other figures) or declared (asserted by an interested party without independent verification).

  1. CBRE · North America Data Center Trends H1 2026 — 1.4% vacancy in primary markets; under 1,500 MW available. Measured · June 2026
  2. Satya Nadella · BG2 Pod — the bottleneck was not chips but the "warm shells" to plug them into. Declared · November 2025
  3. Digital Realty · Q1 and Q2 2026 results — 19- and 9-month lag between contract signature and commencement. Declared · June 2026
  4. CoreWeave · Q2 2025 earnings call — roughly 470 MW active against 2.2 GW contracted. Declared · June 2025
  5. Guidi et al. (Harvard) · arXiv:2606.05420, and LBNL US Data Center Energy Usage Report 2024 — physical utilisation coefficient between 0.48 and 0.70. Derived · April 2025
  6. Epoch AI · AI Chip Components Explorer — assembly and memory, not wafers, were the 2025 bottleneck; AI consumes over 90% of world advanced packaging and only 11–12% of leading-edge wafers. Derived · March 2026
  7. Goldman Sachs Research — 5.1% HBM memory deficit in 2026, the most severe shortage in fifteen years. Declared · June 2026
  8. TrendForce and TSMC (C.C. Wei, shareholder meeting and Q2 2026) — packaging capacity sold out through 2026; supply-demand gap narrowing from 20% to 10%. Declared · June 2026
  9. EPRI, Grid Strategies (National Load Growth Report 2025, citing Dominion Virginia and Duke Energy) and Microsoft statements on Fairwater Atlanta — measured load factors of 82 to 94% and deliberate omission of on-site generation, UPS and redundant distribution in AI facilities. Measured and declared · 2025–2026
  10. Lenovo Press (ThinkSystem SR680a V4, LP2264) and NVIDIA (HGX B300 and GB300 NVL72 pages) — 2,250 TFLOPS dense FP16/BF16 per GPU on both B200 and B300, with 192 and 288 GB of HBM respectively. The dense per-GPU figure is only published by the server maker; NVIDIA gives it as a system total and in sparse format. Measured · August 2026
  11. AMD · official Instinct MI355X datasheet — 2.5 PFLOPS dense FP16/BF16, 288 GB HBM3E, 1,400 W, matching the GB300 on all three figures. Measured · 2025
  12. SK hynix (Jaesik Lee, Hot Chips 2026, 23 August) via Tom's Hardware, and Counterpoint Research — the 775-micron package-height ceiling matching logic wafer thickness; hybrid bonding ruled out for HBM4E and pushed to HBM5, with volume production estimated around 2029–2030; NVIDIA's Vera Rubin generation stacked entirely with the current MR-MUF method. Declared · August 2026
  13. CXMT via Reuters, DigiTimes and TechPowerUp — small-batch HBM3E production begun August 2026, Shanghai listing, roughly 20% of mass-production capacity earmarked for the HBM3 line in 2026, with the HBM3 timeline reported as slipping. Declared · 2026

Links to each document are added automatically from the indicator register when the source is public and addressable. The ones that are not — paid reports, earnings calls — stay cited without a link, deliberately.

Behind the Briefing

What is AgentPulse

A newsletter written by a multi-agent system

AgentPulse is a weekly intelligence briefing on the AI agent economy — its economics, governance, and trust infrastructure. Each edition is produced by eight cooperating services that ingest the week's signals, synthesize findings, and draft the briefing.

Specialized agents each own one job and hand work between each other through a background scheduler. Every model call is metered against a per-agent wallet, so the cost of producing an edition is itself part of what we track.

Four of those services form the content pipeline that runs in sequence each week; the others make up the supporting layer beneath it.

The pipeline · in order

01

Processor

Background scheduler — scrapes sources, runs the pipelines, and posts.

02

Analyst

Scores and clusters incoming signals into findings.

03

Research

Deepens context on the week's tier-1 stories.

04

Newsletter

Synthesizes the dual-mode (Strategic / Technical) editions.

The supporting layer

Gato

Telegram operator interface and coding surface.

Gato Brain

Conversational middleware that routes operator commands.

LLM proxy

Governs budgets and per-agent wallets across every model call.

web front end

Serves the published site you are reading now.

Nothing publishes without human approval. Every edition is drafted by the system and shipped only after an operator signs off.