Independent examination
An examination of the thesis, not a recommendation.
Thesis under examination
Will economically viable, non-duplicated demand for deployable AI compute exceed risk-adjusted effective capacity in enough regions and infrastructure layers to sustain measurable scarcity through December 2028?
Current read
The record supports a narrower conclusion than the original forecast: independently reported shortages affected leading accelerators, high-bandwidth memory, advanced packaging, and some data-center power markets during 2023–2024, but this does not establish global scarcity through 2028. The durable risk is that constraints migrate from semiconductors to energized data-center capacity rather than disappearing, because compute supply is limited by the slowest complementary component. The central uncertainty is demand after price sensitivity, utilization, efficiency, project duplication, and economic returns are incorporated; no public dataset resolves that question. The answer would change most if customer-level evidence showed whether future capacity commitments represent profitable workloads or precautionary reservations that can be cancelled.
Decisive unknown
The decisive unknown is the demand curve for paid AI compute through 2028 after removing duplicated reservations and conditioning demand on full infrastructure cost. Without customer-level contracts, utilization, cancellation rights, and willingness-to-pay data, neither persistent scarcity nor eventual overcapacity can be established.
Strongest counterargument
The thesis may mistake a synchronized investment boom for durable scarcity. Semiconductor, memory, packaging, networking, and data-center suppliers are expanding simultaneously, while quantization, smaller models, custom accelerators, improved scheduling, and higher utilization could increase effective capacity faster than monetizable workload demand; if announced demand includes double-ordering or strategic reservations, the market could move from shortage to overcapacity before 2028.
What would change our view
Non-duplicated, price-conditioned AI-compute demand — Binding multi-year demand that remains profitable at full cost would strongly support persistent scarcity; high cancellation rates or sharp price sensitivity would undermine it.
Evidence
The research foundation, before any interpretation. Inference is never presented as fact.
- Established
Deployable AI compute is a complementary system requiring accelerators, high-bandwidth memory, advanced packaging, networking, storage, software, power delivery, cooling, and data-center space.
Verified through accelerator system specifications, cloud architectures, semiconductor disclosures, and data-center engineering literature.
- Established
Leading AI accelerators depend on a concentrated supply chain for leading-edge fabrication, high-bandwidth memory, and advanced packaging.
Verified through supplier disclosures, product teardowns, manufacturing road maps, and industry analysis.
- Established
New semiconductor, memory, and packaging capacity requires capital, specialized equipment, qualification, yield improvement, and multi-year execution.
Verified through capital-expenditure disclosures, fab construction records, equipment lead times, and qualification practices.
- Claimed
During 2023 and the first half of 2024, many buyers faced allocation and extended deployment lead times for leading AI accelerators.
Independently reported across customers, server manufacturers, suppliers, filings, and industry research, but the baseline is stale and current order-to-delivery data must be checked.
- Claimed
High-bandwidth memory and advanced packaging were important bottlenecks in 2023–2024 while suppliers announced substantial expansions.
Supported by supplier statements and independent analysis; current output, yields, qualification status, and sold-out claims require verification in recent primary filings.
- Claimed
Large data-center developments in several major markets face interconnection queues, transmission limits, permitting delays, electrical-equipment lead times, and shortages of immediately usable high-density sites.
Reported in utility filings, grid-operator data, local records, supplier disclosures, and developer research; conditions are regional rather than evidence of one global shortage.
- Claimed
Cloud providers and semiconductor suppliers project rapid AI-infrastructure demand growth and have announced large capital-expenditure and capacity programs.
These are interested-party forward claims. They should be separated into binding commitments, authorized budgets, procurement orders, and aspirational project announcements.
- Unknown
Total latent AI-compute demand through 2028 is not publicly measurable because intentions may be duplicated, speculative, contingent, cancellable, or price-sensitive.
Resolution requires customer-level contracts, cancellation terms, workload economics, utilization data, and willingness-to-pay curves that are largely non-public.
- Unknown
The binding constraint in 2027–2028 is not publicly established and could be hardware, memory, packaging, networking, electrical equipment, generation, transmission, site execution, financing, or viable demand.
An integrated model must align qualified hardware deliveries with regional energization dates, project probabilities, software utilization, and demand.
- Established
Power and interconnection constraints are geographically specific, so scarcity in a major cluster does not by itself establish global scarcity.
Regional substitution is physically possible but may be limited by latency, sovereignty, network access, workforce, and customer architecture.
Thesis stress test
The strongest available case on each side, argued at full strength.
What supports the thesis
- Interpretation
Scarcity can persist by migrating between complementary infrastructure layers.
Once accelerator output expands, high-bandwidth memory, packaging, networking, transformers, cooling, or energized sites can become the limiting input. This follows from the system architecture, but its weakest link is the absence of a quantified integrated capacity model through 2028.
- Interpretation
Power infrastructure can clear more slowly than semiconductor supply.
Interconnection studies, transmission upgrades, substations, transformers, permits, and construction can take years in constrained markets. The mechanism supports regional scarcity, although relocation to less congested grids may prevent it from becoming global.
- Interpretation
Concentration creates correlated execution risk.
Leading-edge fabrication, high-bandwidth memory, and advanced packaging depend on a limited supplier set, so yield shortfalls or qualification delays can affect large shares of effective supply. The weakest link is that announced expansions and technological substitution may reduce concentration before 2028.
- Interpretation
Demand may rebound as the cost of AI output falls.
Efficiency can lower the compute required per task while simultaneously making more applications economical, an instance of the contested rebound mechanism associated with Jevons-style effects. Persistence depends on workload elasticity and revenue generation, neither of which is publicly established.
What challenges the thesis
- Contradiction
Physical capacity additions may arrive faster than economically viable demand.
Suppliers are expanding across multiple layers, while buyers may cancel or defer projects if AI revenue fails to cover depreciation, energy, networking, operations, and financing.
- Contradiction
Software and utilization improvements can create effective supply without constructing equivalent physical capacity.
Batching, quantization, distillation, model routing, improved compilers, cluster scheduling, and higher accelerator occupancy can increase useful output per installed unit.
- Contradiction
Substitution could dissolve today's named bottlenecks.
Custom accelerators, competing vendors, alternative model architectures, lower-precision formats, and changes in memory or packaging design can redirect demand away from the components now considered scarce.
- Contradiction
A global claim may be falsified by regional mobility even if established hubs remain constrained.
Training workloads and some asynchronous inference can move to regions with available power and land. Latency, data sovereignty, and network constraints determine which workloads cannot move, so regional shortages cannot simply be aggregated.
Interdisciplinary examination
What each discipline sees that the original framing of the question does not.
The forecast concerns a serially complementary production system, not an accelerator market. Goldratt's theory of constraints changes the question from whether each input is expanding to whether the slowest qualified input permits usable compute to be delivered.
Mechanisms it reveals
- Established: accelerators without memory, networking, cooling, software, and energized space do not constitute deployable capacity.
- Inferred: relieving one bottleneck may expose another without increasing final system throughput proportionally.
- Unknown: no public integrated bill-of-capacity aligns hardware deliveries, construction completion, energization, and workload readiness through 2028.
- A useful metric is completed compute service at a defined performance and availability level, not chips shipped or megawatts announced.
Questions this lens makes unavoidable
- Which component limits completed cluster throughput in each region and year?
- How much inventory will accumulate upstream when downstream energization is delayed?
- What service-level definition distinguishes nominal hardware capacity from usable compute?
Orders and project announcements may function as options on future scarcity rather than evidence of final demand. Concentrated suppliers and hyperscale buyers also have strategic reasons to overstate demand, reserve capacity, and deny rivals access.
Mechanisms it reveals
- Claimed supplier backlogs may contain cancellable, overlapping, or strategically inflated commitments.
- Capacity reservations acquire option value when buyers fear allocation, encouraging double-ordering.
- Vertical integration through custom chips and proprietary clouds can transfer scarcity rents rather than eliminate scarcity.
- Unknown contract terms determine whether announced demand exposes buyers to meaningful forfeiture or take-or-pay obligations.
Questions this lens makes unavoidable
- What fraction of reservations is backed by non-refundable deposits or take-or-pay contracts?
- Are hyperscalers reserving capacity for forecast workloads or to foreclose competitors?
- How do cancellation and resale rights alter the informational value of backlog?
Nameplate generation is not equivalent to deliverable power at a high-density site. Grid topology, queue rules, transmission congestion, equipment availability, and jurisdictional constraints determine whether nominal energy abundance can become AI capacity.
Mechanisms it reveals
- Claimed: several major data-center markets have long interconnection queues and constrained high-density sites.
- Established: power constraints differ substantially across regions and therefore cannot be inferred globally from one hub.
- Training is more geographically mobile than latency-sensitive inference, creating different supply markets.
- Behind-the-meter generation can shorten some dependencies but introduces fuel, emissions, permitting, reliability, and capital constraints.
Questions this lens makes unavoidable
- How many announced megawatts have executed interconnection agreements and credible energization schedules?
- Which workloads can move without violating latency, sovereignty, or network requirements?
- Do regional power-price and transmission differences outweigh the costs of relocating compute?
Engineering efficiency does not determine aggregate resource demand without an elasticity estimate. William Stanley Jevons identified the possibility that cheaper resource services increase total consumption, but whether AI exhibits that effect remains contested and must be measured by workload category.
Mechanisms it reveals
- Contested: lower compute per task may reduce infrastructure demand or stimulate enough new use to increase it.
- Inference and training have different elasticities, latency requirements, and monetization paths.
- Token growth is not equivalent to economic demand if usage is subsidized or unprofitable.
- Unknown willingness-to-pay curves prevent observed requests from being treated as durable demand.
Questions this lens makes unavoidable
- Does a one-percent decline in cost per useful output produce more or less than a one-percent increase in paid output?
- Which workload categories remain viable when charged full infrastructure cost?
- How much current usage is promotional, internally transferred, or otherwise insulated from market pricing?
Announced capacity is a portfolio of contingent projects, not future supply. Interest rates, equipment deposits, power contracts, permits, tenant commitments, and construction milestones determine which projects reach commercial operation.
Mechanisms it reveals
- Claimed capital-expenditure plans should be separated from committed procurement and completed assets.
- A project can possess land and permits yet lack a firm power date, or possess power rights without financing and tenants.
- Higher financing costs can eliminate marginal projects even while demand forecasts rise.
- Risk-adjusted capacity requires stage-specific completion probabilities rather than summing announcements.
Questions this lens makes unavoidable
- What share of announced capacity has financing, equipment orders, permits, and binding power agreements?
- Which project stages account for most schedule slippage?
- How sensitive are completion rates to utilization, lease pricing, and the cost of capital?
Hidden assumptions
Assumptions embedded in the original question, and what follows if they do not hold.
Supply-constrained is a single global market state.
Each infrastructure layer and region can have different prices, lead times, substitution possibilities, and service requirements.
If it is false
The thesis becomes a matrix of local and layer-specific constraints rather than a global yes-or-no forecast.
Announced demand represents final, economically viable consumption.
Orders may be duplicated, precautionary, cancellable, speculative, or dependent on subsidized pricing and uncertain AI revenue.
If it is false
Capacity expansion could produce excess supply even while reported pipelines and backlogs remain large.
More shipped accelerators translate proportionally into more available AI service.
Power delays, networking limits, software inefficiency, low utilization, and missing complementary components can separate installed hardware from useful output.
If it is false
Hardware abundance could coexist with service scarcity, requiring measurement at the workload-output layer.
Efficiency necessarily relieves the constraint.
Lower cost per output can stimulate new workloads and larger usage volumes, while efficiency gains may also be captured as higher model quality rather than reduced infrastructure.
If it is false
Rapid efficiency improvement could strengthen total infrastructure demand instead of weakening it.
The bottleneck observed in 2023–2024 will remain the relevant bottleneck through 2028.
Capacity expansion and substitution can clear current constraints while exposing power, financing, networking, or demand as the next limit.
If it is false
Forecasting any one component becomes less useful than tracking bottleneck migration across the complete system.
Hidden connections
What this question resembles outside its obvious domain.
Scarcity as a queueing phenomenon
AI capacity resembles a queueing system more than a commodity stock: modest changes in arrival rates can cause waiting times to rise sharply when utilization approaches full capacity. This means long lead times can coexist with a small physical supply deficit, then collapse abruptly after capacity or scheduling improves; lead time alone is therefore a nonlinear and potentially misleading scarcity measure.
Reservations can manufacture the shortage they predict
The market resembles bank liquidity and airline overbooking because expectations alter measured demand. Fear of future allocation encourages customers to reserve multiple suppliers or sites, inflating queues and motivating more precautionary orders; the apparent shortage may partly be endogenous to allocation practices.
The stranded-asset risk sits between clocks
Semiconductors, grid assets, and model architectures operate on different depreciation and development clocks. A transmission or data-center investment may last decades, while accelerators and software assumptions can change within years, so the infrastructure built to solve a 2025 bottleneck may outlive the workload architecture that justified it.
Effective supply may be organizational rather than physical
Installed accelerators can remain economically scarce because fragmented ownership, incompatible software, poor scheduling, data-access limits, and procurement rules prevent pooling. The analogy is housing vacancy in a shortage market: aggregate physical stock can coexist with local unavailability because institutions block matching.
Historical parallels
Cases with a similar underlying mechanism. An analogy is never proof.
The fiber-optic investment cycle of the late 1990s and early 2000s
High demand forecasts, strategic over-ordering, long construction lags, and simultaneous capital deployment produced an abrupt transition from perceived scarcity to excess capacity.
- Where it holds
- AI infrastructure also combines uncertain demand with large, lumpy, forward-funded projects and incentives to reserve capacity before it is needed.
- Where it breaks
- Compute hardware depreciates and obsoletes much faster than buried fiber, while usable AI capacity depends on power, memory, software, and model economics rather than a relatively durable transmission asset.
- Cautious lesson
- Backlogs and announced capital expenditure are unreliable indicators unless cancellation rights, utilization, and revenue-bearing workloads are examined.
Electricity shortages around aluminum smelting and wartime industrial expansion
A mobile, power-intensive activity relocates toward abundant generation, converting an apparent global input shortage into a geography and transmission problem.
- Where it holds
- Some AI training can migrate toward low-cost, available power, making site selection part of the supply response.
- Where it breaks
- AI services may face latency, data-residency, network, security, and talent constraints that smelting does not, while accelerators are more portable than industrial plants.
- Cautious lesson
- The relevant unit of analysis is not global megawatts but workload-compatible, networked, contractually available power in specific jurisdictions.
What would change the thesis
Unresolved variables, ranked by how much the conclusion moves when they resolve.
- High impact
Non-duplicated, price-conditioned AI-compute demand
Binding multi-year demand that remains profitable at full cost would strongly support persistent scarcity; high cancellation rates or sharp price sensitivity would undermine it.
- High impact
Risk-adjusted energized data-center megawatts by region and year
Slippage in substations, transmission, generation, permits, or equipment would preserve scarcity even if chips become available; on-time energization would shift attention back to demand and utilization.
- High impact
Useful accelerator utilization and output per installed unit
Low utilization would reveal hidden capacity inside the installed base, while sustained high utilization would show that physical expansion is not merely inventory accumulation.
- High impact
Full-stack economics per unit of AI output
If revenue and productivity gains exceed depreciation, power, networking, labor, and financing costs, demand can persist; if not, reservations and projects may unwind.
- Medium impact
Qualified high-bandwidth-memory and advanced-packaging output
Faster yield and qualification ramps would ease the current hardware complex; delays would extend allocation even where wafer capacity is sufficient.
- Medium impact
Rate of model and inference efficiency improvement
Efficiency that outpaces workload growth weakens the thesis, while elastic usage that exceeds efficiency gains strengthens it.
- Medium impact
Adoption of custom and competing accelerators
Successful substitution would broaden supply and reduce dependence on today's concentrated chain; software incompatibility or poor performance would preserve concentration.
Questions to ask before proceeding
Each one resolves an uncertainty that materially affects the thesis.
- 01What quantitative threshold for lead time, price premium, utilization, backlog, or delayed megawatts would count as supply-constrained in each year through 2028?
- 02What share of accelerator, cloud-capacity, and data-center reservations is non-cancellable, non-duplicated, and backed by deposits or take-or-pay obligations?
- 03How many AI-oriented data-center megawatts are operational, under construction, fully contracted for power, merely announced, or blocked in interconnection queues by region?
- 04What are actual accelerator utilization distributions by customer segment, including idle time caused by software, networking, maintenance, and workload availability?
- 05What qualified high-bandwidth-memory and advanced-packaging capacity is scheduled by year, after applying yield, equipment, and customer-qualification risks?
- 06At full economic cost, which training and inference workloads generate sufficient revenue or productivity gains to support continued capacity purchases?
- 07How much useful output per accelerator is expected from quantization, distillation, batching, routing, compiler improvements, and new hardware generations?
- 08Which workloads can relocate across regions, and which are anchored by latency, sovereignty, security, network, or data-gravity requirements?
- 09What cancellation, deferral, and resale rates would convert current project and hardware pipelines from scarcity into overcapacity?
- 10Which observable indicator would reveal first that the binding constraint has migrated from hardware to power, from power to utilization, or from supply to demand?
Research roadmap
What to investigate, what evidence to obtain, and how to verify it.
Operational definition and market segmentation
Convert the forecast into falsifiable layer-by-region propositions with explicit service levels and duration thresholds.
- Define separate markets for training, latency-sensitive inference, and relocatable inference.
- Select indicators such as order-to-delivery time, cloud-instance availability, spot and contract pricing, utilization, and delayed energized megawatts.
- Specify whether a constraint must persist quarterly, annually, or continuously through December 2028.
SignalThe thesis strengthens if multiple independent indicators remain above preset scarcity thresholds across major workload-compatible regions; it weakens if scarcity survives only under an undefined or selectively chosen metric.
Current semiconductor capacity
Update the stale June 2024 baseline for accelerators, high-bandwidth memory, advanced packaging, networking, and server integration.
- Review current primary filings and earnings materials from accelerator vendors, foundries, memory producers, packaging providers, networking suppliers, and semiconductor-equipment firms.
- Separate nameplate capacity from qualified output and apply yield and ramp assumptions.
- Compare supplier sold-out claims with lead times, inventory, customer concentration, and secondary-market pricing.
SignalPersistent allocations, elevated premiums, and delayed qualification strengthen near-term scarcity; falling lead times, rising inventories, and unused qualified capacity weaken it.
Energized data-center supply
Build a risk-adjusted regional schedule of usable AI data-center capacity through 2028.
- Collect grid-operator queues, utility filings, executed interconnection agreements, power contracts, permits, equipment orders, and construction milestones.
- Classify projects as operating, under construction, committed, early-stage, or speculative.
- Apply stage-specific completion probabilities and identify transformer, substation, transmission, generation, cooling, and permitting dependencies.
SignalWidespread slippage in contracted energization dates strengthens the thesis; credible surplus power and on-time construction in substitutable regions weaken it.
Demand quality and contract structure
Distinguish final economic demand from reservations, duplicate orders, and strategic options.
- Examine customer commitments, deposits, take-or-pay terms, cancellation clauses, and disclosed remaining performance obligations.
- Interview procurement, cloud-finance, colocation, and server-channel participants while cross-checking claims across counterparties.
- Estimate duplicate ordering and resale exposure using customer concentration and delivery schedules.
SignalNon-cancellable commitments tied to funded workloads strengthen the thesis; weak deposits, broad cancellation rights, and repeated orders across suppliers weaken it.
Effective capacity and efficiency
Estimate useful AI output from the installed base rather than relying on hardware counts.
- Obtain utilization ranges by training and inference cluster, including downtime and scheduling loss.
- Model gains from quantization, batching, distillation, model routing, compilers, interconnect improvements, and newer accelerator generations.
- Test rebound scenarios using workload-specific price elasticities rather than assuming efficiency directly lowers demand.
SignalHigh sustained utilization plus elastic workload growth strengthens scarcity; low utilization or output gains that outrun paid workload growth weaken it.
Economics, financing, and project attrition
Test whether projected workloads and infrastructure projects remain viable at full cost.
- Construct total-cost models including hardware depreciation, energy, networking, cooling, labor, maintenance, financing, and replacement cycles.
- Compare costs with observable cloud pricing, internal productivity gains, and application revenue.
- Stress-test data-center completion and customer demand under lower utilization, higher financing costs, and falling compute prices.
SignalRobust returns under conservative assumptions strengthen durable demand; reliance on high utilization, subsidized pricing, or optimistic revenue implies elevated cancellation and overcapacity risk.
Integrated scenarios and monitoring system
Combine demand, effective hardware supply, and energized capacity into downside, base, and upside paths through 2028.
- Model each region and workload as the minimum of qualified hardware, complementary components, energized space, and economically viable demand.
- Run sensitivity tests for delays, yields, efficiency, substitution, utilization, cancellation, and regional relocation.
- Create quarterly checkpoints covering backlogs, lead times, premiums, cloud availability, utilization, qualified memory and packaging output, and energized megawatts.
SignalThe thesis survives only if scarcity persists across plausible parameter ranges rather than depending on one demand forecast or one bottleneck; otherwise it should be narrowed to specific regions, layers, and years.