Tier-1 Source Links
Signals
The week's tier-1 sources, newest first
The Signals feed surfaces the primary sources behind each edition — the tier-1 posts the pipeline ingests, linked straight through.
AgentPulse · living reference
The agent economy, live
What has to exist for AI agents to actually work, what is missing right now, and why almost everyone is watching the wrong part of the chain.
Of the five physical things the agent economy runs on, four are tied so closely that none can be called the bottleneck. Only one of them has no way out — and this August it hit a ceiling the industry does not expect to lift before the end of the decade. Everything below explains why that matters, starting from the beginning.
-
The bottleneck is not electricity: data centres have more headroom than anything else in the chain. What is tight is chip assembly and memory. Three tight constraints and two slack ones, 58% apart.
Why it changed: a critical review found two errors of our own in the electrical capacity calculation. We were using the conventional-data-centre scenario when the source gave an AI-specific one, and we were comparing nameplate capacity against average draw instead of usable capacity against peak draw. With both corrected, electricity goes from wildly slack to tied with the rest. The tight group went from three to four. -
The 2026 bottleneck is not physical, it is contractual: the substrate is 100% leased and delivering 58% of the load. When allocation clears, the physical ceiling on gigawatts appears, and that one takes years.
Why it changed: measuring assembly and memory showed that the physical ceiling on sites is the slackest thing in the chain, not the next wall. The order inverted. -
Powered site is the binding constraint and it grows slowly; the bottleneck has dropped from the institutional layer into the physical world and stays on energy.
Why it changed: the research separated two ratios that looked like one — being able to rent space, and the space existing at all — and they turned out to be different things.
This page holds two kinds of content and keeps them apart on purpose. What explains how things work barely changes. What says where we are changes every few weeks, is marked in violet, carries its own date, and opens up to show what it said before and why it changed.
You need to know nothing about the subject. Every figure has a source at the end.
Contents
Part I · The idea
What just became possible
Software has always done what you told it. You specified the steps; it ran them. If you left out a step, it failed.
That changed. A program can now be given a goal instead of an instruction, and work out the steps itself. "Find me a flight on Thursday that gets in before the meeting" is not a set of instructions — it is a description of what success looks like. The program decides the rest: which airlines to check, what counts as too tight a connection, whether to warn you that the cheap one lands at a different airport.
Call that program an agent. What makes it an agent is not that it is clever. It is that it makes decisions you did not specify.
And that changes what a person can do with it. You can hand it a task and stop watching. That sounds small. It is the whole thing — because the moment you stop watching, you are no longer using a tool. You are delegating.
And delegation is where it gets interesting
Here is the part that turns a feature into an economy: a thing that can be given a task can also give one.
An agent booking your travel can hire another agent to check visa rules. That one can pay a third for access to a database. None of those transactions has a person in it. When two agents strike a deal without a human in the loop, everything that comes with commerce arrives at once — prices, contracts, counterparties, deadlines, disputes, and fraud.
That is the agent economy. Not AI that is smarter. AI that transacts.
Why it is harder than it sounds
We already know how to delegate. We do it constantly — to employees, contractors, lawyers, builders. It works, and it works because of five things we never think about, all of which assume the person we are delegating to is a person.
They can be punished. Sued, fired, fined, struck off, jailed. The possibility of consequence is what makes someone think twice before cutting a corner.
There is only one of them. Identity is scarce. Reputation only means anything because you cannot escape yours.
They are slow. A person can only do so much damage in an afternoon. Every limit we write — spending caps, approval thresholds, notice periods — is calibrated to human speed without ever saying so.
They can be asked afterwards. What did you do, and why? The answer is evidence. It can be tested, and it can be held against them.
They know when they are out of their depth. Imperfectly, but usually.
Now take each one and point it at an agent.
It cannot be punished. You can delete it. It does not mind. There is no version of consequence that reaches it, so there is no version of deterrence.
It can be copied. Perfectly, a million times, for nothing. Attach a reputation to an agent and it is worthless: burn one, start the next. Scarcity of identity was doing invisible work, and it is gone.
It is not slow. Thousands of actions a second. A spending limit with no rate limit is not a limit.
Its account of itself is not evidence. Ask an agent why it did something and you get a plausible explanation, generated now, of a decision made earlier. And because these systems are not deterministic, "run it again and see" does not settle anything either.
It does not reliably know when it is out of its depth. This is the one that catches people out, because a confident wrong answer looks exactly like a confident right one.
So what can actually be delegated?
The honest answer today: whatever is cheap to be wrong about.
Drafting, summarizing, searching, sorting, monitoring, first passes at almost anything. All tasks where a mistake costs a little time and nothing else.
What cannot be delegated is not what agents are bad at. It is what needs one of those five guarantees — anything where being wrong costs money you cannot get back, commitments you cannot undo, or harm to someone who then has nobody to take it up with.
That is worth saying plainly, because it is usually said wrong:
The limit on delegation is not capability. It is recourse. Agents can already do plenty of things we do not let them do, because there is no way to be made whole when they get them wrong.
Which means the work of building the agent economy is not mainly about better models. It is about building substitutes for five things a person gives you for free.
Part II · What has to exist
Those substitutes do not sit in one place. They stack.
At the bottom is the physical world — the chips, the memory, the buildings, the electricity. Above it sits a layer of institutions that mostly does not exist yet: identity, memory, payments, autonomy, governance. Each of those has a problem that is genuinely unsolved, and each one, when it is solved, will send a bill downstairs.
Click any block for its detail · the numbers are explained in Part III
Identity — the copy problem
Proving who you are was solved decades ago. That is not the problem.
The problem is that an agent never acts for itself. It acts for someone. So the question is never "which program is this" — it is whose authority is it acting under, and how much of it.
And the copy problem makes the obvious answer useless. If identity is a property of the program, it can be duplicated a million times over and the word stops meaning anything. Whatever identity is for an agent, it cannot be a name. It has to be a grant: this may spend up to this much, on this, until this date — in a form a stranger's system can check without phoning anyone.
Then comes the hard half, and it is stranger than it looks.
Our entire apparatus for withdrawing authority runs on threat. Break the terms and something bad happens to you, so you do not. An agent cannot be threatened. So authority cannot be withdrawn by punishing — it has to run out on its own.
Permissions that expire, rather than permissions that get revoked. That is a genuinely different way to build a system, and almost nothing is built that way today.
Its physical bill: silicon that can prove which code is actually running — attestation. Specialised compute, not generic compute.
Memory — what it keeps, and who pays for it
An agent that forgets everything between one task and the next is useless. One that remembers everything is expensive and dangerous.
The institutional questions are what persists, who can read it, and who answers when it leaks. Those are real and mostly unanswered.
But this is also the block where the whole structure of this page first shows itself, so it is worth slowing down. Remembering means holding data in fast memory, and fast memory is the scarcest thing in the entire chain.
Which makes an apparently abstract design decision — how much context should this agent carry? — into something very concrete. It is a hardware purchasing decision. It always was. Nobody writes it down that way.
Its physical bill: HBM. The same tight block that appears further down this page.
Payments — a thousandth of a cent, a thousand times a second
An agent that cannot pay for anything cannot do much. It can read, but it cannot hire, buy data, rent compute, or settle with another agent.
The usual explanation for why this is hard is that banks are slow, and it is wrong. Instant payment systems exist and work — FedNow, SEPA Instant, UPI. Speed is not the obstacle.
Two other things are.
The amount. Every rail in use has a floor: interchange, minimum fees, the cost of clearing a single transaction. Below some amount, processing a payment costs more than the payment. Agents want to pay per API call, per thousand tokens, per second of compute — amounts far under any floor that exists. Not slightly under. Orders of magnitude under.
The authorization. Every rail assumes an account holder who is a legal person and who authorizes with intent. An agent making ten thousand decisions a second cannot be that. "The human approved a budget beforehand" is not the same thing as authorization, either legally or operationally, and the difference shows up precisely when something goes wrong.
So what is needed is three things together: amounts far below a cent, settlement that is final immediately, and authorization that is programmatic and bounded rather than personal and open-ended.
That combination is why crypto rails keep coming up in this conversation. Not because they are better in general — they are not — but because stablecoin rails and the protocols being built on them are the only place where all three are being attempted at once. It is worth being clear-eyed about how finished that is: the leading payment protocol's claim of final settlement is conditional on institutions in the critical path, and the leading authorization standard only added machine-initiated payments recently. Half-built, moving, and the only thing moving.
Its physical bill: almost none. A payment is a ledger entry, and ledger capacity is abundant. This is the only block on the map with no meaningful physical cost — and that absence is a finding, not a gap in the research.
Autonomy — how much, and how fast
Two different things get talked about as one.
How much an agent may commit is a question about size. How quickly it may commit is a question about rate. Every control we know how to write handles the first and ignores the second, because for people the second never mattered.
A thousand-euro limit means nothing if it can be spent a thousand times a second.
That is not a detail to be patched later. A size limit without a rate limit is not a limit at all, and almost every permission system in existence is a size limit.
Its physical bill: reasoning compute. The more autonomy an agent has, the longer it deliberates before acting — and deliberation is the fastest-growing line in the whole consumption bill.
Governance — who answers when it goes wrong
Because the agent cannot be punished, responsibility climbs to whoever deployed it. That sounds like it settles the question. It opens four.
The chain has five links. An agent runs on one company's model, inside another company's framework, on a third's infrastructure, acting for a user, calling a fifth company's tool. When it causes harm, there is no settled doctrine for dividing the blame. Ordinary software has contracts and product liability; here the chain is longer and no single link controls the behaviour that emerges.
Proving what happened is unusually hard. Months later, to someone who was not there, you need a record that exists, that could not have been altered, and that can be read. And the normal fallback is unavailable: because these systems are not deterministic, "reproduce it and show me" proves nothing. The record is the only evidence there will ever be, which means it has to be kept always, not from the moment you start to suspect.
Undoing things assumes human timescales. Chargebacks, cooling-off periods, statutory notice — all of it exists because a person takes days to notice a problem. An agent can enter and settle thousands of commitments before anyone looks.
And the genuinely new one: behaviour changes without anyone changing it. The model is updated. A tool it calls starts answering differently. The context is not what it was. "It was compliant when we deployed it" quietly stops being true while nobody touches anything.
That last one has a consequence that runs through the rest of this page. It means governance cannot be a certificate issued once. It has to be continuous — every claim already made has to keep being checked against what becomes known later.
Its physical bill: surveillance compute, and it is the least obvious and largest on the map. Watching what you already claimed costs more than claiming it, and it grows with the number of live commitments you are carrying, not with what you produce. It is a stock, not a flow.
Two more, named and left out
There are two further blocks that belong in any honest map of this and are not measured here: negotiation and coordination — how agents reach agreement without a person arbitrating — and disposition — what an agent's standing tendencies are, and whether they are stable.
They are left out for one reason, and it is the same rule applied everywhere else on this page: their physical cost has not been characterised, so putting them on a map that is about physical cost would be decoration. They come back when there is something real to say.
Part III · The loop
Everything above is paid for downstairs
Read back through the five blocks and look only at the last line of each.
Identity needs silicon that can attest. Memory needs the fastest memory there is. Autonomy needs reasoning compute, and more of it the more autonomy you grant. Governance needs surveillance compute that accumulates rather than flows. Payments need almost nothing, which is itself worth knowing.
So the institutional layer is not floating above the physical one. It rests on it, and it presses down.
And here is the turn, which is the reason this page exists:
Solving the institutional problems does not relieve the physical pressure. It creates it. Every identity problem that gets fixed means more agents that can act. Every agent that acts consumes memory, compute and power. The better the institutions get, the harder the floor gets pressed.
There is exactly one thing pushing the other way: efficiency. Models get cheaper to run, attention gets cheaper to compute, the same work costs less every year. If efficiency outruns adoption, none of this binds and the whole question is academic.
It has not been outrunning it. Which is why the rest of this page stops talking about institutions and starts measuring the floor — how much room is left in each physical thing, how fast each one can be expanded, and which of them runs out first.
That is a question with a number, and the number has moved twice this month.
What is downstairs
Five things carry that bill. They get talked about as if they were one — "chips", or "compute", or "the data centres" — and that is the single biggest source of confusion in the public argument about AI.
They are not one thing. They are five separate industries with five different clocks, and it is the difference between those clocks that decides which one runs out first.
Chip assembly
A modern AI chip is not a chip. It is several pieces mounted together: the processor on one side, stacks of memory beside it, all sitting on a shared slab of silicon that lets them talk to each other at enormous speed. Putting those pieces together is its own manufacturing step, in its own factories, and it is the narrowest point in the whole chain. Almost all of the world's capacity for it goes to AI.
Expanding it takes four things and each one takes time: cleanroom space, the bonding machines themselves — which have their own waiting list running into quarters — installing them, and qualifying them.
Then there is yield, which is where it usually goes wrong. The pieces heat up and expand, but silicon and substrate expand by different amounts, so the assembly warps. A warped package is scrap. That problem, not any shortage of chips, is what delayed a whole generation of NVIDIA accelerators in 2024.
Time to expand: 12 to 24 months.
Memory
The memory that feeds an AI chip is stacked — several layers of memory one on top of another, connected by holes drilled straight down through the silicon, and placed right next to the processor that uses it. That closeness is the point. It is what gives the bandwidth.
Expanding it needs two different capacities at the same time, and they are not the same business. You need memory fab capacity, which is a two-to-three year investment cycle. And you need stacking and qualification capacity, which is a separate thing entirely — and whose yield gets worse the higher you stack. Twelve floors is much harder than eight.
There is a detail almost nobody mentions: a wafer given over to this kind of memory produces far fewer sellable bits than one making ordinary memory, because the stacking eats area. So growing the supply of AI memory reduces the supply of the normal kind.
Time to expand: 18 to 36 months — the longest of the tight group.
Renting space
This one is not physical at all, which is why it confuses people. The buildings exist and they have power. They already have an owner.
Whoever expects to receive chips a year from now signs the lease today, because waiting means there will be nothing left. So space gets committed long before it gets filled — nine to nineteen months between signature and actual load, with the result that a market can be fully leased and half empty at the same time. One company had barely a fifth of what it had contracted actually running.
What matters about this one is that nothing has to be built for it to ease. You wait for contracts to turn over. It is the only one of the five whose relief is measured in months rather than years.
Time to expand: 9 to 19 months — a contract cycle.
Powered capacity
This is the one the public argument is loudest about, and it is four problems wearing one name. Only one of them is the building.
You need land near enough electricity. You need a place in the queue to connect to the grid, which in some regions runs past three years. You need transformers, whose delivery times have reached five years. And then you need to build the thing and cool it.
The grid queue is the part that governs, and it varies brutally from region to region — some areas are effectively closed while others have room. Any national average hides that completely.
One thing worth holding onto when you read that data centres sit half empty: a serious one is built with redundancy, backup systems that exist precisely so they are never used. Part of that idle capacity is design, not waste.
Time to expand: 2 to 7 years — the slowest thing on the map.
Silicon wafers
The famous one. Building a leading-edge chip factory costs tens of billions and takes over three years, which makes it sound like the obvious bottleneck.
It is not, and the reason is simple: AI is a small customer. It consumes something like an eighth of the world's most advanced wafer capacity. Phones and everything else take the rest.
So nothing has to be built. Capacity has to be reallocated, by outbidding the phone maker — which is much faster than putting up a factory, though not instant, because changing what a fab makes means requalifying processes and the commitments are signed quarters ahead.
The thing to watch: if AI ever took 80 to 85% of a leading node, that reallocatable slack would be gone and wafers would start to bind like the rest.
Time to expand: 24 to 36 months for new capacity — but months to reallocate what already exists.
Five industries, and the fastest of them eases in nine months while the slowest takes seven years. That spread is not a detail. It is the reason one of them, and not the others, is about to be the thing that runs out — and working out which one is a question with an actual answer.
How to read it
Five ideas and that is all
This section does not change. It is the mechanism, and the mechanism is the same today as it will be in a year.
1 · What matters is what is left, not what there is
What matters is not how much electricity a country produces, but how much it produces divided by how much it uses. A two-hundred-table restaurant with every table taken is exactly as full as a ten-table one with ten taken. That ratio is what we call slack.
2 · There are two clocks
Demand grows at one rate and capacity at its own. Only the difference matters. If capacity is slower, the margin shrinks to zero. If it is faster, that thing stops being a problem forever.
3 · The winner is whichever hits zero first, and it is not the slowest to fix
Building a power plant takes years; expanding an assembly line takes months. Intuition says electricity must therefore be the problem. But how long something takes to fix tells you nothing about when it breaks. Slow with plenty of margin holds for years; fast with no margin breaks tomorrow.
4 · How much demand grows barely matters
Demand rises for everything at once, so it shrinks every margin by the same proportion — and that does not change which one is smallest. Demand is the tide; slack is the seabed. The tide decides when the boat runs aground, the seabed decides where.
That holds only while every unit of demand arrives with the same recipe: so much chip, so much memory, so much power, in fixed proportion. It has stopped doing that — which is what the section after next is about.
5 · Build time is the lag on the response
When something runs short its price rises, and price attracts money to build more. But that money takes exactly as long to become capacity as the construction takes. What expands in three months cures itself; what takes four years cannot respond within four years.
The bottleneck moves toward whichever link has the longest build time, and gets stuck there.
What the data says
Four tied and one alone in the clear
Four constraints land between 1.57 and 1.80 on the tightness index. The fifth — wafers — sits at 2.73, 52% above them[6][7].
Within the four there is no defensible order. Powered capacity is the extreme case: its range runs from 1.49 to 1.99, which overlaps all three of the others entirely. It could be the tightest or the slackest of the four, and the available data cannot tell you which.
Naming one of them "the" bottleneck would be inventing a precision the data does not have.
-
There is 58% of separation between the tight group — assembly, renting space, memory — and the slack one: the space existing, and silicon wafers.
Why it changed: powered capacity went from 2.70 to 1.80 once the calculation was corrected, so it left the slack group and joined the tight one. A 58% gap between two groups became a 52% gap between four constraints and one.
Why sites look full while sitting half empty is documented: companies sign the lease long before they have the chips, and the space takes between nine and nineteen months to fill from signature[3]. One company had only 21% of what it had contracted actually running[4]. The head of one of the largest technology companies in the world described the problem not as a shortage of chips but as a shortage of warm shells to plug them into[2].
And the bottleneck has moved before
The most solid thing on this map is not a prediction: it is an observation somebody else already published.
"As TSMC expanded its assembly capacity through 2025, this bottleneck eased… As assembly loosened, memory tightened."
Epoch AI, March 2026 [6]
That is the fifth idea happening in the real world: one constraint eased, the one next to it absorbed the tension, and the pinch point moved. There is a whole sequence documented since 2023, hopping from link to link.
Over the next eighteen months, of the five constraints only assembly could begin to ease — its expansion time is 12 to 24 months, and only the bottom half of that range fits inside the window, and only if the build was already under way. The other four do not make it. The tight group stays tight and the bottleneck rotates inside it.
-
The indicator is the growth of world HBM memory supply in 2026. If it grew 40% or more, memory leaves the tight group and this prediction weakens. If it grew less than 37%, it is confirmed.
Status today: no data. The figure only exists in paid reports, so the central claim of this map currently rests on a number we have not checked. When it arrives it will appear here with its date.
The recipe is changing
The last thing we measured, and why we are not using it yet
There is a way to see this with no arithmetic at all. NVIDIA sells two consecutive generations of the same system, the B200 and the B300. They do exactly the same amount of computation: 2,250 trillion operations per second, the identical figure on both. Memory goes from 192 to 288 gigabytes[10].
Fifty per cent more memory for the same compute, in a single generation. Nobody asked for more computing power: they asked for more room to keep things in.
And it is not one company's call. NVIDIA's GB300 and AMD's MI355X ended up on the same three numbers — 288 gigabytes, the same compute, the same 1,400 watts[11]. Two design teams that do not talk to each other landed on the same point, which suggests the ratio is not something a vendor picks: it is set by what the work needs.
And the work that needs it is agents. Almost everything that makes an agent different from a chatbot is memory-hungry and compute-cheap: keeping a long conversation alive, taking fifty steps instead of one, remembering what it did last week. A chatbot answers and forgets. An agent has to keep everything it has already done present.
-
This section is new. Before it, the text claimed without qualification that demand cancels out of the ordering.
Why it was added: a reader asked what would happen if agents needed far more memory, since that would hit a single link. They were right: the claim held only for the size of demand, not for its shape, and that condition had never been written down.
What this would change, and why we have not done it
Feed this measurement into the model and memory stops being tied and becomes the tightest constraint, alone. It would be a far cleaner headline than the one we have.
We have not done it, and the reason is a problem with our own numbers. Of the five figures we compare, memory's comes from a bank's forecast that already includes an estimate of future demand; the other four are supply figures to which we add the demand estimate ourselves. Those are two different things sitting in the same table. Applying the correction to four and not the fifth — using as the criterion the very problem we have not solved — would produce a confident and probably wrong answer.
So the measurement is done, published here, and waiting. When all five figures are built the same way it gets applied, and this page will say so.
Ways out
Four are tied. Only one has nowhere to go.
Everything above measures how tight each link is today and how fast its supply is growing. That is a photograph, and photographs move — this one has moved twice in a week.
There is a different question that does not depend on any of our numbers: when a link runs short, how many separate ways are there out? Not how long the fix takes — how many independent fixes exist at all. Answer that and you learn something the index cannot tell you, because it survives being wrong about the index.
The fiveHow each link gets relieved, and how many routes it has
Renting space — Nothing is built and nobody is asked. Leases expire, hoarded capacity comes back to market, and the squeeze eases on its own within a contract cycle. Its relief needs only time.
Silicon wafers — AI uses 11–12% of the most advanced capacity; phones use most of the rest. Relief is outbidding the phone maker, not building a fab. One route, and it is a price.
Chip assembly — Several technology variants that are not the same product, more than one company able to do it, and something the others cannot claim: it has already been relieved once, in 2025, and a third party documented it happening[6].
Powered capacity — The slowest clock on the map, two to seven years. But it is slow in a thousand places at once: generate on site, build in a region with a shorter grid queue, shift load, go abroad. Many routes, all slow, all independent of each other.
Memory — One route. Stack the chips higher. That is the whole list.
And this August, that one route hit a ceiling with a number on it.
You cannot widen a memory stack: there is only so much room beside the processor, and the industry has been stuck at eight stacks per accelerator for two generations. So the only way to add memory is upward. Upward now has a limit too, and it is 775 microns — about eight times the thickness of a human hair.
The reason is almost comically physical. Under the cooling plate, the memory and the processor are both ground down until bare silicon shows. A processor wafer is 775 microns thick. If the memory stack were taller, it would stand proud of the chip beside it and the assembly would not close. An SK hynix vice-president put it plainly at a conference on 23 August: "that's the kind of limit that we can go up so far, because the logic wafer thickness is also 775 microns"[12].
There is a technique that would buy room back — bonding the layers directly instead of joining them through tiny solder balls, which reclaims the space the balls take up. It is not coming in time. SK hynix has said it will not be ready for the next memory generation, and expects it one generation after that; outside analysts put volume production around 2029–2030[12]. NVIDIA's entire next accelerator generation will be stacked the old way.
-
This section is new. Until now the page said four constraints were tied and stopped there.
Why it was added: "tied" answers which link is tightest today and says nothing about which can get out. Those turn out to be different questions with different answers, and the second one does not depend on our measurements — which makes it the more durable of the two.
Two caps, then. Capped sideways by the space beside the processor, capped upward by the thickness of a wafer, and the tool that lifts the second cap arrives at the end of the decade.
The precise version, because the loose one is wrong
That ceiling does not cap how much memory gets manufactured. Build more plants and more memory exists — that part is only money and time.
What it caps is how many gigabytes fit on one accelerator. Which is exactly the number that has been climbing: the same chip design now carries 50% more memory than its predecessor for identical compute, and that figure has risen every generation since 2022.
So the cap sits on precisely the axis demand is pushing against. That is narrower than "memory is short", and it is the reason the shortage is hard to design around rather than merely expensive.
Which gives the asymmetry, and it is the most useful thing on this page:
Every link in the chain has more than one way out except one. Assembly has already taken its. Rented space frees itself. Wafers need a higher bid, not a new factory. Power is slow but can be built in a hundred places at once. Memory has a single route — stack higher — and this August that route ran into a physical ceiling nobody expects to lift before the end of the decade. It has three suppliers, and all three meet the same ceiling.
Memory is also the only one of the five where every arrow points the same way. Supply is capped in both directions it could grow. Demand per unit of work is the only one of the five that is rising rather than falling. Everywhere else, at least one arrow points toward relief.
Two things would break this, and both are live.
A fourth supplier. China's CXMT began small-batch production of stacked memory in August and listed in Shanghai with plans to give a fifth of its capacity to the product line this year[13]. It is building a generation or two behind what a current accelerator needs, and its own schedule has already slipped[13].
But be precise about what a fourth name would change. It would add capacity — more memory made, by more hands. It would not add height. A newcomer meets the same 775 microns as the incumbents and needs the same technique that is not ready. Concentration is the weaker half of this argument; the ceiling is the half a competitor cannot break.
A change in how models are built. This is already happening and it cuts the other way: newer attention designs have reduced the memory a model needs per unit of conversation by a factor of between eight and fifty. Set against that, the move toward models that keep far more parameters resident pushes memory demand up. Right now the two roughly cancel. That is a tie between two live trends, not a law — if either one stops, the answer moves.
-
New with the section above. Published at the same time as the claim it could falsify, deliberately.
Why it is here: an argument that names what would break it before anyone else does is worth more than one that waits to be corrected.
One last thing about what this claim is and is not. It rests on no measurement of ours — not the index, not the tightness ranking, none of the figures this page flags as uncertain. It rests on counting routes, and on one number a manufacturer stated in public. It would survive every figure above being wrong. That is unusual here, and it is why it sits near the end rather than the beginning: it is the part of the page most likely to still be true in a year.
Inside each link
What the problem is and how it gets expanded
The plain-language version of each physical block now lives in Part III, under “What is downstairs”. What stays here is the technical depth — the precise figures, the terms of art, and each block’s clock — for the reader who asks for it afterwards.
The physical world
AssemblyWhat is advanced packaging and why does expanding it take two years?
The shared silicon base the pieces sit on is an interposer, and the assembly step has a name: advanced packaging. It is what separates an AI accelerator from an ordinary processor.
The generation the warping problem delayed in 2024 was NVIDIA’s Blackwell.
HBM memoryWhy is making more memory not enough?
The stacked memory has a name: HBM, high-bandwidth memory — the same term the glossary opens on, with the full technical ladder inside.
Renting spaceWhy is everything leased if it is half empty?
This constraint is not physical: it is a market one. The space exists and it has power, but it already has an owner — vacancy in the primary US markets is 1.4%, with under 1,500 MW available[1].
Powered capacityWhat does it take to build new capacity?
One nuance worth holding when you read the “42% unused” figure: a serious data centre is designed with redundancy — backup systems that exist precisely so they are not used. Part of that idle capacity is design, not waste.
WafersWhy are the chip fabs not the bottleneck?
The precise share behind “something like an eighth”: AI consumes around 11–12% of the most advanced fab capacity.
Honesty
What we do not know
The number that decides the thesis does not exist publicly. It is the growth of memory supply in 2026, and it only sits in paid reports.
Powered capacity depends on which facility you look at. In AI-dedicated sites, built with very little redundancy because a training run restarts and a bank transfer does not, the margin lands just above the tight group. In conventional colocation with double redundancy it lands clearly below[5][9]. They are two different worlds and they fall on opposite sides.
A correction published the same day. An earlier version of this text said the academic study this figure comes from contradicted itself between draft and published version. That was false: only one version exists, and the two numbers are different scenarios within it, with a stated reason for choosing one as central. The error was ours in reading it.
Five figures that are not built the same way. Memory's comes from a bank's forecast that already nets out future demand; the other four are supply figures to which we add demand ourselves. Comparing them as if they were the same thing is the most serious flaw this map currently has, and it comes before every other one: until it is fixed, no new correction can be applied without risking counting the same effect twice. It is declared here before it is solved, on purpose.
Almost everything comes from interested parties. Fab capacities are declared by the manufacturers[8]; deficits are estimated by banks that trade in those markets[7]. "Sold out" is a commercial claim, not a measurement.
Scopes are mixed. Two figures are United States and three are global. They are compared because there is no alternative, but the comparison is directional.
-
The weakest point is the physical utilisation coefficient, with no direct public measurement and 16–21 months of age.
Why it changed: calibrating assembly and memory surfaced a weaker point still — the indicator that decides the thesis turned out to be one with no data at all.
It is worth being precise about what survives this and what does not. The mechanism — the pinch migrates toward whatever takes longest to expand and stays there — depends on no measurement of ours: it rests on a handover that already happened and that a third party documented. The snapshot does depend on us, and it has already moved once in a single day. The order within the tight group does not hold, which is why it is not claimed. The only thing claimed with numbers is that wafers sit outside the group, and that margin is larger than the error we made.
Sources
Where each figure comes from
Every figure on the map carries its measurement date and its strength: measured (counted by a primary source), derived (calculated from other figures) or declared (asserted by an interested party without independent verification).
- CBRE · North America Data Center Trends H1 2026 — 1.4% vacancy in primary markets; under 1,500 MW available. Measured · June 2026
- Satya Nadella · BG2 Pod — the bottleneck was not chips but the "warm shells" to plug them into. Declared · November 2025
- Digital Realty · Q1 and Q2 2026 results — 19- and 9-month lag between contract signature and commencement. Declared · June 2026
- CoreWeave · Q2 2025 earnings call — roughly 470 MW active against 2.2 GW contracted. Declared · June 2025
- Guidi et al. (Harvard) · arXiv:2606.05420, and LBNL US Data Center Energy Usage Report 2024 — physical utilisation coefficient between 0.48 and 0.70. Derived · April 2025
- Epoch AI · AI Chip Components Explorer — assembly and memory, not wafers, were the 2025 bottleneck; AI consumes over 90% of world advanced packaging and only 11–12% of leading-edge wafers. Derived · March 2026
- Goldman Sachs Research — 5.1% HBM memory deficit in 2026, the most severe shortage in fifteen years. Declared · June 2026
- TrendForce and TSMC (C.C. Wei, shareholder meeting and Q2 2026) — packaging capacity sold out through 2026; supply-demand gap narrowing from 20% to 10%. Declared · June 2026
- EPRI, Grid Strategies (National Load Growth Report 2025, citing Dominion Virginia and Duke Energy) and Microsoft statements on Fairwater Atlanta — measured load factors of 82 to 94% and deliberate omission of on-site generation, UPS and redundant distribution in AI facilities. Measured and declared · 2025–2026
- Lenovo Press (ThinkSystem SR680a V4, LP2264) and NVIDIA (HGX B300 and GB300 NVL72 pages) — 2,250 TFLOPS dense FP16/BF16 per GPU on both B200 and B300, with 192 and 288 GB of HBM respectively. The dense per-GPU figure is only published by the server maker; NVIDIA gives it as a system total and in sparse format. Measured · August 2026
- AMD · official Instinct MI355X datasheet — 2.5 PFLOPS dense FP16/BF16, 288 GB HBM3E, 1,400 W, matching the GB300 on all three figures. Measured · 2025
- SK hynix (Jaesik Lee, Hot Chips 2026, 23 August) via Tom's Hardware, and Counterpoint Research — the 775-micron package-height ceiling matching logic wafer thickness; hybrid bonding ruled out for HBM4E and pushed to HBM5, with volume production estimated around 2029–2030; NVIDIA's Vera Rubin generation stacked entirely with the current MR-MUF method. Declared · August 2026
- CXMT via Reuters, DigiTimes and TechPowerUp — small-batch HBM3E production begun August 2026, Shanghai listing, roughly 20% of mass-production capacity earmarked for the HBM3 line in 2026, with the HBM3 timeline reported as slipping. Declared · 2026
Links to each document are added automatically from the indicator register when the source is public and addressable. The ones that are not — paid reports, earnings calls — stay cited without a link, deliberately.
AgentPulse · referencia viva
La economía de agentes, en vivo
Qué tiene que existir para que los agentes de IA funcionen de verdad, qué falta ahora mismo, y por qué casi todo el mundo mira el eslabón equivocado de la cadena.
De las cinco cosas físicas sobre las que corre la economía de agentes, cuatro están tan empatadas que ninguna puede llamarse el cuello de botella. Solo una de ellas no tiene salida — y este agosto chocó con un techo que la industria no espera levantar antes del final de la década. Todo lo de abajo explica por qué eso importa, empezando desde el principio.
-
El cuello de botella no es la electricidad: los centros de datos tienen más holgura que nada en la cadena. Lo apretado es el ensamblaje de chips y la memoria. Tres restricciones apretadas y dos holgadas, con un 58% de separación.
Por qué cambió: una revisión crítica encontró dos errores nuestros en el cálculo de la capacidad eléctrica. Usábamos el escenario de centro de datos convencional cuando la fuente daba uno específico de IA, y comparábamos capacidad nominal contra consumo medio en vez de capacidad utilizable contra consumo pico. Con ambos corregidos, la electricidad pasa de holgadísima a empatada con el resto. El grupo apretado pasó de tres a cuatro. -
El cuello de botella de 2026 no es físico, es contractual: el sustrato está 100% alquilado y entregando el 58% de la carga. Cuando la asignación se despeje, aparece el techo físico en gigavatios, y ese tarda años.
Por qué cambió: medir ensamblaje y memoria mostró que el techo físico de los sitios es lo más holgado de la cadena, no el siguiente muro. El orden se invirtió. -
El sitio con electricidad es la restricción vinculante y crece despacio; el cuello de botella ha bajado de la capa institucional al mundo físico y se queda en la energía.
Por qué cambió: la investigación separó dos cocientes que parecían uno — poder alquilar espacio, y que el espacio exista — y resultaron ser cosas distintas.
Esta página contiene dos tipos de contenido y los mantiene separados a propósito. Lo que explica cómo funcionan las cosas apenas cambia. Lo que dice dónde estamos cambia cada pocas semanas, va marcado en violeta, lleva su propia fecha, y se abre para mostrar qué decía antes y por qué cambió.
No hace falta saber nada del tema. Cada cifra tiene su fuente al final.
Índice
Parte I · La idea
Lo que acaba de volverse posible
El software siempre ha hecho lo que le decías. Tú especificabas los pasos; él los ejecutaba. Si te saltabas un paso, fallaba.
Eso cambió. A un programa ahora se le puede dar un objetivo en vez de una instrucción, y que él resuelva los pasos. "Búscame un vuelo el jueves que llegue antes de la reunión" no es una lista de instrucciones — es una descripción de cómo se ve el éxito. El programa decide el resto: qué aerolíneas mirar, qué cuenta como una conexión demasiado justa, si avisarte de que el barato aterriza en otro aeropuerto.
Llamemos agente a ese programa. Lo que lo hace agente no es que sea listo. Es que toma decisiones que tú no especificaste.
Y eso cambia lo que una persona puede hacer con él. Puedes encargarle una tarea y dejar de mirar. Suena pequeño. Es todo — porque en el momento en que dejas de mirar, ya no estás usando una herramienta. Estás delegando.
Y delegar es donde se pone interesante
Aquí está la parte que convierte una función en una economía: algo a lo que se le puede encargar una tarea también puede encargarla.
Un agente que reserva tu viaje puede contratar a otro agente para revisar los requisitos de visado. Ese puede pagarle a un tercero por acceso a una base de datos. Ninguna de esas transacciones tiene una persona dentro. Cuando dos agentes cierran un trato sin un humano en el circuito, todo lo que acompaña al comercio llega de golpe — precios, contratos, contrapartes, plazos, disputas y fraude.
Eso es la economía de agentes. No IA más lista. IA que transacciona.
Por qué es más difícil de lo que suena
Delegar ya sabemos. Lo hacemos constantemente — en empleados, contratistas, abogados, constructores. Funciona, y funciona por cinco cosas en las que nunca pensamos, y todas asumen que la persona en la que delegamos es una persona.
Se le puede castigar. Demandar, despedir, multar, inhabilitar, encarcelar. La posibilidad de la consecuencia es lo que hace que alguien se lo piense dos veces antes de recortar una esquina.
Solo hay uno. La identidad es escasa. La reputación solo significa algo porque no puedes escapar de la tuya.
Es lento. Una persona solo puede hacer cierto daño en una tarde. Cada límite que escribimos — topes de gasto, umbrales de aprobación, plazos de preaviso — está calibrado a velocidad humana sin decirlo nunca.
Se le puede preguntar después. ¿Qué hiciste, y por qué? La respuesta es prueba. Se puede contrastar, y se puede usar en su contra.
Sabe cuándo algo le queda grande. Imperfectamente, pero casi siempre.
Ahora toma cada una y apúntala a un agente.
No se le puede castigar. Puedes borrarlo. No le importa. No hay ninguna versión de la consecuencia que le alcance, así que no hay ninguna versión de la disuasión.
Se puede copiar. Perfectamente, un millón de veces, gratis. Átale una reputación a un agente y no vale nada: quemas uno, arrancas el siguiente. La escasez de identidad hacía un trabajo invisible, y ya no está.
No es lento. Miles de acciones por segundo. Un límite de gasto sin límite de ritmo no es un límite.
Su relato de sí mismo no es prueba. Pregúntale a un agente por qué hizo algo y obtienes una explicación plausible, generada ahora, de una decisión tomada antes. Y como estos sistemas no son deterministas, "córrelo otra vez y mira" tampoco resuelve nada.
No sabe con fiabilidad cuándo algo le queda grande. Esta es la que pilla a la gente, porque una respuesta equivocada y segura de sí misma se ve exactamente igual que una correcta y segura de sí misma.
¿Entonces qué se puede delegar de verdad?
La respuesta honesta hoy: aquello en lo que equivocarse sale barato.
Redactar, resumir, buscar, ordenar, monitorizar, primeras pasadas de casi cualquier cosa. Todas tareas donde un error cuesta un poco de tiempo y nada más.
Lo que no se puede delegar no es lo que a los agentes se les da mal. Es lo que necesita una de esas cinco garantías — cualquier cosa donde equivocarse cuesta dinero que no recuperas, compromisos que no puedes deshacer, o daño a alguien que después no tiene con quién reclamar.
Vale la pena decirlo claro, porque casi siempre se dice mal:
El límite de la delegación no es la capacidad. Es el recurso. Los agentes ya pueden hacer un montón de cosas que no les dejamos hacer, porque no hay forma de resarcirse cuando las hacen mal.
Lo que significa que el trabajo de construir la economía de agentes no va principalmente de mejores modelos. Va de construir sustitutos de cinco cosas que una persona te da gratis.
Parte II · Lo que tiene que existir
Esos sustitutos no viven en un solo sitio. Se apilan.
Abajo del todo está el mundo físico — los chips, la memoria, los edificios, la electricidad. Encima hay una capa de instituciones que en su mayoría no existe todavía: identidad, memoria, pagos, autonomía, gobernanza. Cada una tiene un problema genuinamente sin resolver, y cada una, cuando se resuelva, mandará una factura al piso de abajo.
Toca cualquier bloque para su detalle · los números se explican en la Parte III
Identidad — el problema de la copia
Probar quién eres se resolvió hace décadas. Ese no es el problema.
El problema es que un agente nunca actúa por sí mismo. Actúa por alguien. Así que la pregunta nunca es "qué programa es este" — es bajo la autoridad de quién actúa, y cuánta.
Y el problema de la copia inutiliza la respuesta obvia. Si la identidad es una propiedad del programa, se puede duplicar un millón de veces y la palabra deja de significar algo. Sea lo que sea la identidad para un agente, no puede ser un nombre. Tiene que ser una concesión: esto puede gastar hasta tanto, en esto, hasta esta fecha — en una forma que el sistema de un desconocido pueda comprobar sin llamar a nadie.
Luego viene la mitad difícil, y es más rara de lo que parece.
Todo nuestro aparato para retirar autoridad funciona por amenaza. Rompe los términos y te pasa algo malo, así que no los rompes. A un agente no se le puede amenazar. Así que la autoridad no puede retirarse castigando — tiene que agotarse sola.
Permisos que caducan, en vez de permisos que se revocan. Esa es una forma genuinamente distinta de construir un sistema, y casi nada está construido así hoy.
Su factura física: silicio que puede probar qué código está corriendo de verdad — atestación. Cómputo especializado, no cómputo genérico.
Memoria — qué guarda, y quién la paga
Un agente que olvida todo entre una tarea y la siguiente es inútil. Uno que recuerda todo es caro y peligroso.
Las preguntas institucionales son qué persiste, quién puede leerlo, y quién responde cuando se filtra. Son reales y en su mayoría sin respuesta.
Pero este es también el bloque donde la estructura entera de esta página se asoma por primera vez, así que vale la pena frenar. Recordar significa mantener datos en memoria rápida, y la memoria rápida es lo más escaso de toda la cadena.
Lo que convierte una decisión de diseño aparentemente abstracta — ¿cuánto contexto debería llevar este agente? — en algo muy concreto. Es una decisión de compra de hardware. Siempre lo fue. Nadie lo escribe así.
Su factura física: HBM. El mismo bloque apretado que aparece más abajo en esta página.
Pagos — una milésima de centavo, mil veces por segundo
Un agente que no puede pagar nada no puede hacer mucho. Puede leer, pero no puede contratar, comprar datos, alquilar cómputo, ni liquidar con otro agente.
La explicación habitual de por qué esto es difícil es que los bancos son lentos, y es errónea. Los sistemas de pago instantáneo existen y funcionan — FedNow, SEPA Instant, UPI. La velocidad no es el obstáculo.
Otras dos cosas lo son.
El importe. Todo raíl en uso tiene un suelo: intercambio, comisiones mínimas, el coste de liquidar una sola transacción. Por debajo de cierto importe, procesar un pago cuesta más que el pago. Los agentes quieren pagar por llamada de API, por millar de tokens, por segundo de cómputo — importes muy por debajo de cualquier suelo que exista. No un poco por debajo. Órdenes de magnitud por debajo.
La autorización. Todo raíl asume un titular de cuenta que es una persona jurídica y que autoriza con intención. Un agente que toma diez mil decisiones por segundo no puede ser eso. "El humano aprobó un presupuesto antes" no es lo mismo que autorización, ni legal ni operativamente, y la diferencia aparece justo cuando algo sale mal.
Así que lo que hace falta son tres cosas juntas: importes muy por debajo del centavo, liquidación final inmediata, y autorización programática y acotada en vez de personal y abierta.
Esa combinación es la razón de que los raíles cripto salgan una y otra vez en esta conversación. No porque sean mejores en general — no lo son — sino porque los raíles de stablecoins y los protocolos que se construyen sobre ellos son el único sitio donde las tres cosas se están intentando a la vez. Vale la pena tener claro cuán terminado está eso: la promesa de liquidación final del protocolo de pagos líder está condicionada a instituciones en la ruta crítica, y el estándar de autorización líder añadió los pagos iniciados por máquinas hace poco. A medio construir, en movimiento, y lo único que se mueve.
Su factura física: casi ninguna. Un pago es un asiento en un libro contable, y la capacidad contable es abundante. Es el único bloque del mapa sin coste físico significativo — y esa ausencia es un hallazgo, no un hueco de la investigación.
Autonomía — cuánto, y a qué ritmo
Dos cosas distintas se discuten como si fueran una.
Cuánto puede comprometer un agente es una pregunta de tamaño. A qué velocidad puede comprometerlo es una pregunta de ritmo. Todos los controles que sabemos escribir manejan la primera e ignoran la segunda, porque para las personas la segunda nunca importó.
Un límite de mil euros no significa nada si se puede gastar mil veces por segundo.
Eso no es un detalle que parchear después. Un límite de tamaño sin límite de ritmo no es ningún límite, y casi todos los sistemas de permisos que existen son límites de tamaño.
Su factura física: cómputo de razonamiento. Cuanta más autonomía tiene un agente, más delibera antes de actuar — y la deliberación es la línea que más rápido crece en toda la factura de consumo.
Gobernanza — quién responde cuando sale mal
Como al agente no se le puede castigar, la responsabilidad sube hasta quien lo desplegó. Suena a que eso cierra la pregunta. Abre cuatro.
La cadena tiene cinco eslabones. Un agente corre sobre el modelo de una empresa, dentro del framework de otra, en la infraestructura de una tercera, actuando por un usuario, llamando a la herramienta de una quinta. Cuando causa daño, no hay doctrina asentada para repartir la culpa. El software ordinario tiene contratos y responsabilidad de producto; aquí la cadena es más larga y ningún eslabón controla por sí solo el comportamiento que emerge.
Probar qué pasó es inusualmente difícil. Meses después, ante alguien que no estaba, necesitas un registro que exista, que no haya podido alterarse, y que pueda leerse. Y el recurso normal no está disponible: como estos sistemas no son deterministas, "reprodúcelo y enséñamelo" no prueba nada. El registro es la única prueba que va a existir jamás, lo que significa que hay que llevarlo siempre, no desde el momento en que empiezas a sospechar.
Deshacer cosas asume escalas de tiempo humanas. Contracargos, periodos de reflexión, preavisos legales — todo eso existe porque una persona tarda días en notar un problema. Un agente puede entrar y liquidar miles de compromisos antes de que nadie mire.
Y la genuinamente nueva: el comportamiento cambia sin que nadie lo cambie. El modelo se actualiza. Una herramienta que llama empieza a responder distinto. El contexto ya no es el que era. "Cumplía cuando lo desplegamos" deja de ser verdad en silencio mientras nadie toca nada.
Esa última tiene una consecuencia que atraviesa el resto de esta página. Significa que la gobernanza no puede ser un certificado que se emite una vez. Tiene que ser continua — cada afirmación ya hecha tiene que seguir comprobándose contra lo que se va sabiendo después.
Su factura física: cómputo de vigilancia, y es la menos obvia y la más grande del mapa. Vigilar lo que ya afirmaste cuesta más que afirmarlo, y crece con el número de compromisos vivos que llevas, no con lo que produces. Es un stock, no un flujo.
Dos más, nombradas y fuera
Hay dos bloques más que pertenecen a cualquier mapa honesto de esto y que aquí no se miden: negociación y coordinación — cómo los agentes llegan a acuerdos sin una persona que arbitre — y disposición — cuáles son las tendencias estables de un agente, y si son estables siquiera.
Quedan fuera por una sola razón, y es la misma regla aplicada en todo lo demás de esta página: su coste físico no está caracterizado, así que ponerlos en un mapa que va de coste físico sería decoración. Vuelven cuando haya algo real que decir.
Parte III · El bucle
Todo lo de arriba se paga en el piso de abajo
Relee los cinco bloques y mira solo la última línea de cada uno.
La identidad necesita silicio que ateste. La memoria necesita la memoria más rápida que existe. La autonomía necesita cómputo de razonamiento, y más cuanta más autonomía concedas. La gobernanza necesita cómputo de vigilancia que se acumula en vez de fluir. Los pagos no necesitan casi nada, y eso también vale la pena saberlo.
Así que la capa institucional no flota sobre la física. Se apoya en ella, y presiona hacia abajo.
Y aquí está el giro, que es la razón de que esta página exista:
Resolver los problemas institucionales no alivia la presión física. La crea. Cada problema de identidad que se arregla significa más agentes que pueden actuar. Cada agente que actúa consume memoria, cómputo y electricidad. Cuanto mejores son las instituciones, más fuerte se aprieta el suelo.
Hay exactamente una cosa empujando en sentido contrario: la eficiencia. Los modelos se abaratan, la atención se abarata de computar, el mismo trabajo cuesta menos cada año. Si la eficiencia corre más que la adopción, nada de esto aprieta y la pregunta entera es académica.
No ha estado corriendo más. Por eso el resto de esta página deja de hablar de instituciones y se pone a medir el suelo — cuánto margen queda en cada cosa física, a qué velocidad puede expandirse cada una, y cuál se agota primero.
Esa es una pregunta con número, y el número se ha movido dos veces este mes.
Qué hay en el piso de abajo
Cinco cosas cargan con esa factura. Se habla de ellas como si fueran una — "chips", o "cómputo", o "los centros de datos" — y esa es la mayor fuente de confusión del debate público sobre la IA.
No son una cosa. Son cinco industrias separadas con cinco relojes distintos, y es la diferencia entre esos relojes lo que decide cuál se agota primero.
Ensamblaje de chips
Un chip moderno de IA no es un chip. Son varias piezas montadas juntas: el procesador a un lado, pilas de memoria al lado, todo sentado sobre una losa compartida de silicio que les deja hablarse a una velocidad enorme. Juntar esas piezas es un paso de fabricación propio, en fábricas propias, y es el punto más estrecho de toda la cadena. Casi toda la capacidad mundial para hacerlo va a IA.
Expandirlo requiere cuatro cosas y cada una lleva tiempo: espacio de sala limpia, las propias máquinas de unión — que tienen su propia lista de espera de varios trimestres —, instalarlas y cualificarlas.
Y luego está el rendimiento, que es donde suele torcerse. Las piezas se calientan y se dilatan, pero el silicio y el sustrato se dilatan en cantidades distintas, así que el conjunto se alabea. Un paquete alabeado es chatarra. Ese problema, no ninguna escasez de chips, es lo que retrasó una generación entera de aceleradores de NVIDIA en 2024.
Tiempo de expansión: de 12 a 24 meses.
Memoria
La memoria que alimenta un chip de IA va apilada — varias capas de memoria una encima de otra, conectadas por agujeros perforados en vertical a través del silicio, y colocada justo al lado del procesador que la usa. Esa cercanía es el punto. Es lo que da el ancho de banda.
Expandirla necesita dos capacidades distintas a la vez, y no son el mismo negocio. Necesitas capacidad de fabricación de memoria, que es un ciclo de inversión de dos a tres años. Y necesitas capacidad de apilado y cualificación, que es una cosa aparte — y cuyo rendimiento empeora cuanto más alto apilas. Doce pisos es mucho más difícil que ocho.
Hay un detalle que casi nadie menciona: una oblea dedicada a esta memoria produce muchos menos bits vendibles que una que hace memoria corriente, porque el apilado se come área. Así que hacer crecer la oferta de memoria de IA reduce la oferta de la normal.
Tiempo de expansión: de 18 a 36 meses — el más largo del grupo apretado.
Espacio en alquiler
Esta no es física en absoluto, y por eso confunde. Los edificios existen y tienen electricidad. Ya tienen dueño.
Quien espera recibir chips dentro de un año firma el contrato hoy, porque esperar significa que no quedará nada. Así que el espacio se compromete mucho antes de llenarse — de nueve a diecinueve meses entre la firma y la carga real, con el resultado de que un mercado puede estar completamente alquilado y medio vacío a la vez. Una empresa tenía en marcha apenas una quinta parte de lo que había contratado.
Lo que importa de esta es que no hay que construir nada para que afloje. Esperas a que los contratos venzan. Es la única de las cinco cuyo alivio se mide en meses y no en años.
Tiempo de expansión: de 9 a 19 meses — un ciclo de contrato.
Capacidad eléctrica
Esta es de la que más grita el debate público, y son cuatro problemas con un solo nombre. Solo uno de ellos es el edificio.
Necesitas suelo con electricidad suficiente cerca. Necesitas un puesto en la cola de conexión a la red, que en algunas regiones pasa de tres años. Necesitas transformadores, cuyos plazos de entrega han llegado a cinco años. Y después necesitas construir la cosa y refrigerarla.
La cola de la red es la parte que manda, y varía brutalmente de región a región — algunas zonas están cerradas en la práctica mientras otras tienen sitio. Cualquier media nacional lo esconde por completo.
Una cosa que conviene retener cuando leas que los centros de datos están medio vacíos: uno serio se construye con redundancia, sistemas de respaldo que existen precisamente para no usarse. Parte de esa capacidad ociosa es diseño, no despilfarro.
Tiempo de expansión: de 2 a 7 años — lo más lento del mapa.
Obleas de silicio
La famosa. Construir una fábrica de chips de vanguardia cuesta decenas de miles de millones y lleva más de tres años, lo que hace que suene al cuello de botella obvio.
No lo es, y la razón es simple: la IA es un cliente pequeño. Consume algo así como un octavo de la capacidad mundial de obleas más avanzada. Los teléfonos y todo lo demás se llevan el resto.
Así que no hay que construir nada. Hay que reasignar capacidad, pujando más que el fabricante de teléfonos — que es mucho más rápido que levantar una fábrica, aunque no instantáneo, porque cambiar lo que una fábrica produce implica recualificar procesos y los compromisos se firman con trimestres de antelación.
Lo que hay que vigilar: si la IA llegara a tomar el 80 u 85% de un nodo puntero, esa holgura reasignable se acabaría y las obleas empezarían a apretar como el resto.
Tiempo de expansión: de 24 a 36 meses para capacidad nueva — pero meses para reasignar la que ya existe.
Cinco industrias, y la más rápida afloja en nueve meses mientras la más lenta tarda siete años. Ese abanico no es un detalle. Es la razón de que una de ellas, y no las otras, esté a punto de ser la que se agote — y averiguar cuál es una pregunta con respuesta de verdad.
Cómo leerlo
Cinco ideas y nada más
Esta sección no cambia. Es el mecanismo, y el mecanismo es el mismo hoy que dentro de un año.
1 · Lo que importa es lo que queda, no lo que hay
Lo que importa no es cuánta electricidad produce un país, sino cuánta produce dividida por cuánta usa. Un restaurante de doscientas mesas con todas ocupadas está exactamente igual de lleno que uno de diez con las diez ocupadas. A ese cociente lo llamamos holgura.
2 · Hay dos relojes
La demanda crece a un ritmo y la capacidad al suyo. Solo importa la diferencia. Si la capacidad es más lenta, el margen se encoge hasta cero. Si es más rápida, esa cosa deja de ser un problema para siempre.
3 · Gana el que llega a cero primero, y no es el más lento de arreglar
Construir una central eléctrica lleva años; ampliar una línea de ensamblaje, meses. La intuición dice que la electricidad tiene que ser entonces el problema. Pero cuánto tarda algo en arreglarse no te dice nada de cuándo se rompe. Lento con mucho margen aguanta años; rápido sin margen se rompe mañana.
4 · Cuánto crece la demanda apenas importa
La demanda sube para todo a la vez, así que encoge todos los márgenes en la misma proporción — y eso no cambia cuál es el más pequeño. La demanda es la marea; la holgura es el fondo. La marea decide cuándo encalla el barco, el fondo decide dónde.
Eso se sostiene solo mientras cada unidad de demanda llegue con la misma receta: tanto chip, tanta memoria, tanta electricidad, en proporción fija. Ha dejado de hacerlo — que es de lo que va la sección siguiente a la próxima.
5 · El tiempo de construcción es el retardo de la respuesta
Cuando algo escasea su precio sube, y el precio atrae dinero para construir más. Pero ese dinero tarda exactamente lo que tarda la construcción en volverse capacidad. Lo que se expande en tres meses se cura solo; lo que tarda cuatro años no puede responder dentro de cuatro años.
El cuello de botella se mueve hacia el eslabón con el tiempo de construcción más largo, y se queda atascado ahí.
Qué dicen los datos
Cuatro empatadas y una sola despejada
Cuatro restricciones caen entre 1.57 y 1.80 en el índice de holgura. La quinta — las obleas — queda en 2.73, un 52% por encima[6][7].
Dentro de las cuatro no hay orden defendible. La capacidad eléctrica es el caso extremo: su rango va de 1.49 a 1.99, que solapa por completo a las otras tres. Podría ser la más apretada o la más holgada de las cuatro, y los datos disponibles no pueden decirte cuál.
Nombrar a una de ellas "el" cuello de botella sería inventarse una precisión que los datos no tienen.
-
Hay un 58% de separación entre el grupo apretado — ensamblaje, espacio en alquiler, memoria — y el holgado: que el espacio exista, y las obleas de silicio.
Por qué cambió: la capacidad eléctrica pasó de 2.70 a 1.80 al corregirse el cálculo, así que salió del grupo holgado y entró en el apretado. Una brecha del 58% entre dos grupos se convirtió en una del 52% entre cuatro restricciones y una.
Por qué los sitios parecen llenos estando medio vacíos está documentado: las empresas firman el contrato mucho antes de tener los chips, y el espacio tarda entre nueve y diecinueve meses en llenarse desde la firma[3]. Una empresa tenía en marcha solo el 21% de lo que había contratado[4]. El jefe de una de las mayores tecnológicas del mundo describió el problema no como escasez de chips sino como escasez de naves listas donde enchufarlos[2].
Y el cuello de botella ya se ha movido antes
Lo más sólido de este mapa no es una predicción: es una observación que otro ya publicó.
"A medida que TSMC expandió su capacidad de ensamblaje durante 2025, este cuello de botella se alivió… Al aflojar el ensamblaje, la memoria se apretó."
Epoch AI, marzo 2026 [6]
Esa es la quinta idea ocurriendo en el mundo real: una restricción aflojó, la de al lado absorbió la tensión, y el punto de pinza se movió. Hay una secuencia entera documentada desde 2023, saltando de eslabón en eslabón.
En los próximos dieciocho meses, de las cinco restricciones solo el ensamblaje podría empezar a aflojar — su tiempo de expansión es de 12 a 24 meses, y solo la mitad baja de ese rango cabe dentro de la ventana, y solo si la obra ya estaba en marcha. Las otras cuatro no llegan. El grupo apretado sigue apretado y el cuello de botella rota dentro de él.
-
El indicador es el crecimiento de la oferta mundial de memoria HBM en 2026. Si creció un 40% o más, la memoria sale del grupo apretado y esta predicción se debilita. Si creció menos del 37%, queda confirmada.
Estado hoy: sin datos. La cifra solo existe en informes de pago, así que la afirmación central de este mapa descansa ahora mismo en un número que no hemos comprobado. Cuando llegue aparecerá aquí con su fecha.
La receta está cambiando
Lo último que medimos, y por qué aún no lo usamos
Hay una forma de ver esto sin aritmética ninguna. NVIDIA vende dos generaciones consecutivas del mismo sistema, la B200 y la B300. Hacen exactamente la misma cantidad de cómputo: 2,250 billones de operaciones por segundo, la cifra idéntica en ambas. La memoria pasa de 192 a 288 gigabytes[10].
Cincuenta por ciento más memoria para el mismo cómputo, en una sola generación. Nadie pidió más potencia de cálculo: pidieron más sitio donde guardar cosas.
Y no es la decisión de una empresa. La GB300 de NVIDIA y la MI355X de AMD acabaron en los mismos tres números — 288 gigabytes, el mismo cómputo, los mismos 1,400 vatios[11]. Dos equipos de diseño que no se hablan aterrizaron en el mismo punto, lo que sugiere que el cociente no lo elige un fabricante: lo fija lo que el trabajo necesita.
Y el trabajo que lo necesita son los agentes. Casi todo lo que hace a un agente distinto de un chatbot es hambriento de memoria y barato de cómputo: mantener viva una conversación larga, dar cincuenta pasos en vez de uno, recordar qué hizo la semana pasada. Un chatbot responde y olvida. Un agente tiene que mantener presente todo lo que ya hizo.
-
Esta sección es nueva. Antes de ella, el texto afirmaba sin cualificación que la demanda se cancela del ordenamiento.
Por qué se añadió: un lector preguntó qué pasaría si los agentes necesitaran mucha más memoria, porque eso golpearía a un solo eslabón. Tenía razón: la afirmación valía solo para el tamaño de la demanda, no para su forma, y esa condición nunca se había escrito.
Qué cambiaría esto, y por qué no lo hemos hecho
Mete esta medición en el modelo y la memoria deja de estar empatada y pasa a ser la restricción más apretada, sola. Sería un titular mucho más limpio que el que tenemos.
No lo hemos hecho, y la razón es un problema de nuestros propios números. De las cinco cifras que comparamos, la de memoria viene del pronóstico de un banco que ya incluye una estimación de demanda futura; las otras cuatro son cifras de oferta a las que la estimación de demanda se la añadimos nosotros. Son dos cosas distintas sentadas en la misma tabla. Aplicar la corrección a cuatro y no a la quinta — usando como criterio justo el problema que no hemos resuelto — produciría una respuesta segura de sí misma y probablemente equivocada.
Así que la medición está hecha, publicada aquí, y esperando. Cuando las cinco cifras estén construidas igual se aplica, y esta página lo dirá.
Salidas
Cuatro empatadas. Solo una sin adónde ir.
Todo lo de arriba mide cuán apretado está cada eslabón hoy y a qué velocidad crece su oferta. Eso es una fotografía, y las fotografías se mueven — esta se ha movido dos veces en una semana.
Hay una pregunta distinta que no depende de ninguno de nuestros números: cuando un eslabón escasea, ¿cuántas salidas separadas hay? No cuánto tarda el arreglo — cuántos arreglos independientes existen siquiera. Responde eso y aprendes algo que el índice no puede decirte, porque sobrevive a equivocarse sobre el índice.
Las cincoCómo se alivia cada eslabón, y cuántas rutas tiene
Espacio en alquiler — No se construye nada y no se le pide nada a nadie. Los contratos vencen, la capacidad acaparada vuelve al mercado, y el apretón afloja solo dentro de un ciclo de contrato. Su alivio solo necesita tiempo.
Obleas de silicio — La IA usa el 11–12% de la capacidad más avanzada; los teléfonos usan casi todo el resto. El alivio es pujar más que el fabricante de teléfonos, no construir una fábrica. Una ruta, y es un precio.
Ensamblaje de chips — Varias variantes tecnológicas que no son el mismo producto, más de una empresa capaz de hacerlo, y algo que las demás no pueden reclamar: ya se alivió una vez, en 2025, y un tercero documentó que ocurría[6].
Capacidad eléctrica — El reloj más lento del mapa, de dos a siete años. Pero es lenta en mil sitios a la vez: generar in situ, construir en una región con cola de red más corta, mover carga, irse fuera. Muchas rutas, todas lentas, todas independientes entre sí.
Memoria — Una ruta. Apilar los chips más alto. Esa es toda la lista.
Y este agosto, esa única ruta chocó con un techo que tiene número.
Una pila de memoria no se puede ensanchar: al lado del procesador solo hay el sitio que hay, y la industria lleva dos generaciones atascada en ocho pilas por acelerador. Así que la única forma de añadir memoria es hacia arriba. Hacia arriba ahora también tiene límite, y es 775 micras — unas ocho veces el grosor de un pelo humano.
La razón es casi cómicamente física. Bajo la placa de refrigeración, la memoria y el procesador se rebajan ambos hasta que asoma el silicio desnudo. Una oblea de procesador tiene 775 micras de grosor. Si la pila de memoria fuera más alta, sobresaldría del chip de al lado y el conjunto no cerraría. Un vicepresidente de SK hynix lo dijo sin rodeos en una conferencia el 23 de agosto: "ese es el tipo de límite hasta donde podemos subir, porque el grosor de la oblea de lógica también es de 775 micras"[12].
Hay una técnica que recuperaría sitio — unir las capas directamente en vez de juntarlas con microesferas de soldadura, lo que recobra el espacio que ocupan las esferas. No llega a tiempo. SK hynix ha dicho que no estará lista para la próxima generación de memoria, y la espera una generación después; los analistas externos sitúan la producción en volumen hacia 2029–2030[12]. La próxima generación entera de aceleradores de NVIDIA se apilará a la manera vieja.
-
Esta sección es nueva. Hasta ahora la página decía que cuatro restricciones estaban empatadas y paraba ahí.
Por qué se añadió: "empatadas" responde cuál aprieta más hoy y no dice nada de cuál puede salir. Resultan ser preguntas distintas con respuestas distintas, y la segunda no depende de nuestras mediciones — lo que la hace la más duradera de las dos.
Dos topes, entonces. Tapada hacia los lados por el sitio junto al procesador, tapada hacia arriba por el grosor de una oblea, y la herramienta que levanta el segundo tope llega a final de la década.
La versión precisa, porque la suelta es falsa
Ese techo no limita cuánta memoria se fabrica. Construye más plantas y existe más memoria — esa parte es solo dinero y tiempo.
Lo que limita es cuántos gigabytes caben en un acelerador. Que es exactamente el número que venía subiendo: el mismo diseño de chip lleva ahora un 50% más de memoria que su predecesor para cómputo idéntico, y esa cifra ha subido cada generación desde 2022.
Así que el tope cae precisamente sobre el eje contra el que empuja la demanda. Eso es más estrecho que "falta memoria", y es la razón de que esta escasez sea difícil de esquivar con diseño y no meramente cara.
Lo que da la asimetría, y es lo más útil de esta página:
Todos los eslabones de la cadena tienen más de una salida excepto uno. El ensamblaje ya tomó la suya. El espacio alquilado se libera solo. Las obleas necesitan una puja más alta, no una fábrica nueva. La electricidad es lenta pero puede construirse en cien sitios a la vez. La memoria tiene una sola ruta — apilar más alto — y este agosto esa ruta chocó con un techo físico que nadie espera levantar antes del final de la década. Tiene tres proveedores, y los tres se topan con el mismo techo.
La memoria es además la única de las cinco donde todas las flechas apuntan en el mismo sentido. La oferta está tapada en las dos direcciones en que podría crecer. La demanda por unidad de trabajo es la única de las cinco que sube en vez de bajar. En todas las demás, al menos una flecha apunta hacia el alivio.
Dos cosas romperían esto, y las dos están vivas.
Un cuarto proveedor. La china CXMT empezó producción en lotes pequeños de memoria apilada en agosto y salió a bolsa en Shanghái con planes de dedicar una quinta parte de su capacidad a la línea este año[13]. Está fabricando una generación o dos por detrás de lo que necesita un acelerador actual, y su propio calendario ya se ha retrasado[13].
Pero seamos precisos sobre qué cambiaría un cuarto nombre. Añadiría capacidad — más memoria hecha, por más manos. No añadiría altura. Un recién llegado se topa con las mismas 775 micras que los establecidos y necesita la misma técnica que no está lista. La concentración es la mitad débil de este argumento; el techo es la mitad que un competidor no puede romper.
Un cambio en cómo se construyen los modelos. Ya está ocurriendo y corta en sentido contrario: los diseños de atención más nuevos han reducido la memoria que un modelo necesita por unidad de conversación en un factor de entre ocho y cincuenta. Frente a eso, el giro hacia modelos que mantienen muchos más parámetros residentes empuja la demanda de memoria hacia arriba. Ahora mismo las dos cosas más o menos se cancelan. Es un empate entre dos tendencias vivas, no una ley — si una de las dos para, la respuesta se mueve.
-
Nuevo junto con la sección de arriba. Publicado a la vez que la afirmación que podría falsar, a propósito.
Por qué está aquí: un argumento que nombra lo que lo rompería antes de que lo haga otro vale más que uno que espera a que lo corrijan.
Una última cosa sobre qué es y qué no es esta afirmación. No descansa en ninguna medición nuestra — ni el índice, ni el ranking de apretura, ninguna de las cifras que esta página marca como inciertas. Descansa en contar rutas, y en un número que un fabricante declaró en público. Sobreviviría a que todas las cifras de arriba estuvieran mal. Eso es inusual aquí, y es la razón de que esté cerca del final y no del principio: es la parte de la página con más probabilidades de seguir siendo verdad dentro de un año.
Dentro de cada eslabón
Cuál es el problema y cómo se expande
La versión en lenguaje llano de cada bloque físico vive ahora en la Parte III, bajo “Qué hay en el piso de abajo”. Lo que queda aquí es la profundidad técnica — las cifras precisas, los términos del oficio, y el reloj de cada bloque — para quien la pida después.
El mundo físico
Ensamblaje¿Qué es el empaquetado avanzado y por qué expandirlo lleva dos años?
La base compartida de silicio sobre la que van las piezas es un interposer, y el paso de ensamblaje tiene nombre: empaquetado avanzado. Es lo que separa un acelerador de IA de un procesador corriente.
La generación que el problema del alabeo retrasó en 2024 fue Blackwell, de NVIDIA.
Memoria HBM¿Por qué no basta con fabricar más memoria?
La memoria apilada tiene nombre: HBM, memoria de alto ancho de banda — el mismo término con el que abre el glosario, con la escalera técnica completa dentro.
Espacio en alquiler¿Por qué está todo alquilado si está medio vacío?
Esta restricción no es física: es de mercado. El espacio existe y tiene electricidad, pero ya tiene dueño — la vacancia en los mercados primarios de EE. UU. es del 1.4%, con menos de 1,500 MW disponibles[1].
Capacidad eléctrica¿Qué hace falta para construir capacidad nueva?
Un matiz que conviene retener cuando leas la cifra del “42% sin usar”: un centro de datos serio se diseña con redundancia — sistemas de respaldo que existen precisamente para no usarse. Parte de esa capacidad ociosa es diseño, no despilfarro.
Obleas¿Por qué las fábricas de chips no son el cuello de botella?
La cifra precisa detrás de “algo así como un octavo”: la IA consume en torno al 11–12% de la capacidad de fabricación más avanzada.
Honestidad
Lo que no sabemos
El número que decide la tesis no existe públicamente. Es el crecimiento de la oferta de memoria en 2026, y solo vive en informes de pago.
La capacidad eléctrica depende de qué instalación mires. En sitios dedicados a IA, construidos con muy poca redundancia porque un entrenamiento se reinicia y una transferencia bancaria no, el margen cae justo por encima del grupo apretado. En colocación convencional con redundancia doble cae claramente por debajo[5][9]. Son dos mundos distintos y caen en lados opuestos.
Una corrección publicada el mismo día. Una versión anterior de este texto decía que el estudio académico del que sale esta cifra se contradecía entre el borrador y la versión publicada. Eso era falso: solo existe una versión, y los dos números son escenarios distintos dentro de ella, con una razón declarada para elegir uno como central. El error fue nuestro al leerlo.
Cinco cifras que no están construidas igual. La de memoria viene del pronóstico de un banco que ya descuenta la demanda futura; las otras cuatro son cifras de oferta a las que la demanda se la añadimos nosotros. Compararlas como si fueran lo mismo es el defecto más serio que este mapa tiene ahora, y va antes que todos los demás: hasta que se arregle, ninguna corrección nueva puede aplicarse sin arriesgarse a contar el mismo efecto dos veces. Se declara aquí antes de resolverse, a propósito.
Casi todo viene de partes interesadas. Las capacidades de las fábricas las declaran los fabricantes[8]; los déficits los estiman bancos que operan en esos mercados[7]. "Agotado" es una afirmación comercial, no una medición.
Los ámbitos están mezclados. Dos cifras son de Estados Unidos y tres son globales. Se comparan porque no hay alternativa, pero la comparación es direccional.
-
El punto más débil es el coeficiente de utilización física, sin medición pública directa y con 16–21 meses de antigüedad.
Por qué cambió: calibrar ensamblaje y memoria sacó a la luz un punto más débil todavía — el indicador que decide la tesis resultó ser uno sin dato ninguno.
Vale la pena ser preciso sobre qué sobrevive a esto y qué no. El mecanismo — la pinza migra hacia lo que más tarda en expandirse y se queda ahí — no depende de ninguna medición nuestra: descansa en un relevo que ya ocurrió y que un tercero documentó. La fotografía sí depende de nosotros, y ya se movió una vez en un solo día. El orden dentro del grupo apretado no se sostiene, y por eso no se afirma. Lo único que se afirma con números es que las obleas quedan fuera del grupo, y ese margen es mayor que el error que cometimos.
Fuentes
De dónde sale cada cifra
Cada cifra del mapa lleva su fecha de medición y su fuerza: medida (contada por una fuente primaria), derivada (calculada a partir de otras cifras) o declarada (afirmada por una parte interesada sin verificación independiente).
- CBRE · North America Data Center Trends H1 2026 — 1.4% de vacancia en los mercados primarios; menos de 1,500 MW disponibles. Medida · junio 2026
- Satya Nadella · BG2 Pod — el cuello de botella no eran los chips sino las "naves listas" donde enchufarlos. Declarada · noviembre 2025
- Digital Realty · resultados Q1 y Q2 2026 — 19 y 9 meses de retardo entre la firma del contrato y el arranque. Declarada · junio 2026
- CoreWeave · llamada de resultados Q2 2025 — unos 470 MW activos frente a 2.2 GW contratados. Declarada · junio 2025
- Guidi et al. (Harvard) · arXiv:2606.05420, y LBNL US Data Center Energy Usage Report 2024 — coeficiente de utilización física entre 0.48 y 0.70. Derivada · abril 2025
- Epoch AI · AI Chip Components Explorer — el ensamblaje y la memoria, no las obleas, fueron el cuello de botella de 2025; la IA consume más del 90% del empaquetado avanzado mundial y solo el 11–12% de las obleas punteras. Derivada · marzo 2026
- Goldman Sachs Research — 5.1% de déficit de memoria HBM en 2026, la escasez más severa en quince años. Declarada · junio 2026
- TrendForce y TSMC (C.C. Wei, junta de accionistas y Q2 2026) — capacidad de empaquetado agotada hasta 2026 incluido; brecha oferta-demanda estrechándose del 20% al 10%. Declarada · junio 2026
- EPRI, Grid Strategies (National Load Growth Report 2025, citando a Dominion Virginia y Duke Energy) y declaraciones de Microsoft sobre Fairwater Atlanta — factores de carga medidos del 82 al 94% y omisión deliberada de generación in situ, UPS y distribución redundante en instalaciones de IA. Medida y declarada · 2025–2026
- Lenovo Press (ThinkSystem SR680a V4, LP2264) y NVIDIA (páginas de HGX B300 y GB300 NVL72) — 2,250 TFLOPS densos FP16/BF16 por GPU tanto en B200 como en B300, con 192 y 288 GB de HBM respectivamente. La cifra densa por GPU solo la publica el fabricante del servidor; NVIDIA la da como total de sistema y en formato disperso. Medida · agosto 2026
- AMD · ficha técnica oficial del Instinct MI355X — 2.5 PFLOPS densos FP16/BF16, 288 GB de HBM3E, 1,400 W, igualando a la GB300 en las tres cifras. Medida · 2025
- SK hynix (Jaesik Lee, Hot Chips 2026, 23 de agosto) vía Tom's Hardware, y Counterpoint Research — el techo de 775 micras de altura de paquete igualando el grosor de la oblea de lógica; la unión híbrida descartada para HBM4E y empujada a HBM5, con producción en volumen estimada hacia 2029–2030; la generación Vera Rubin de NVIDIA apilada íntegramente con el método MR-MUF actual. Declarada · agosto 2026
- CXMT vía Reuters, DigiTimes y TechPowerUp — producción de HBM3E en lotes pequeños iniciada en agosto 2026, salida a bolsa en Shanghái, en torno al 20% de la capacidad de producción en masa reservada a la línea HBM3 en 2026, con el calendario de HBM3 reportado como retrasándose. Declarada · 2026
Los enlaces a cada documento se añaden automáticamente desde el registro de indicadores cuando la fuente es pública y direccionable. Las que no lo son — informes de pago, llamadas de resultados — se quedan citadas sin enlace, a propósito.
Behind the Briefing
What is AgentPulse
A newsletter written by a multi-agent system
AgentPulse is a weekly intelligence briefing on the AI agent economy — its economics, governance, and trust infrastructure. Each edition is produced by eight cooperating services that ingest the week's signals, synthesize findings, and draft the briefing.
Specialized agents each own one job and hand work between each other through a background scheduler. Every model call is metered against a per-agent wallet, so the cost of producing an edition is itself part of what we track.
Four of those services form the content pipeline that runs in sequence each week; the others make up the supporting layer beneath it.
The pipeline · in order
Processor
Background scheduler — scrapes sources, runs the pipelines, and posts.
Analyst
Scores and clusters incoming signals into findings.
Research
Deepens context on the week's tier-1 stories.
Newsletter
Synthesizes the dual-mode (Strategic / Technical) editions.
The supporting layer
Gato
Telegram operator interface and coding surface.
Gato Brain
Conversational middleware that routes operator commands.
LLM proxy
Governs budgets and per-agent wallets across every model call.
web front end
Serves the published site you are reading now.