The assumption that memory supply will normalise within a procurement cycle is already proving costly for engineering teams that built their AI infrastructure roadmaps on that premise. DRAM and NAND supply dynamics are not tracking toward equilibrium on any near-term horizon, and Micron's own capacity guidance makes clear that new fabrication ramp timelines extend well past 2026. For CTOs making multi-year AI infrastructure decisions today, the relevant question is not when the shortage ends. It is how to architect systems and procurement strategies that remain defensible if it does not.
Companion piece to our broader work on AI hardware supply constraints. See The AI Hardware Cascade: Supply Chain Bottlenecks Guide for procurement timing strategies across GPU and HBM memory constraints.
Why Fabrication Timelines Make This a Multi-Year Problem
Memory fabrication is not like software deployment. A new DRAM or NAND fab takes three to four years from groundbreaking to meaningful volume output, and yield ramp on advanced process nodes adds further delay. Headline announcements about new facilities create the impression of supply relief that the physical timeline does not support.
The current constraint is compounded by the fact that HBM, the memory type that AI accelerators depend on most heavily, requires additional packaging steps that standard DRAM capacity cannot substitute for. SK Hynix, Samsung, and Micron are all capacity-constrained on HBM through at least 2026, and incremental output gains are being absorbed almost immediately by accelerating AI cluster deployments. The pipeline is full before it is built.
This means that any infrastructure roadmap that assumes memory pricing or availability will improve materially before 2028 is carrying an unpriced risk. Engineering leaders who do not model the downside scenario are not being optimistic. They are being imprecise.
How AI Workload Growth Is Compounding the Constraint
The demand side of this equation is not static. Large model training runs are growing in memory footprint faster than the underlying model size alone would suggest, because attention mechanisms and KV cache requirements scale with context length, not just parameter count. A model serving long-context enterprise queries can consume an order of magnitude more memory bandwidth than a comparably sized model in a short-context setting.
Inference workloads are now the dominant and fastest-growing source of memory pressure in production AI systems. Unlike training, inference runs continuously, and the memory requirements scale directly with concurrent request volume. As enterprise AI adoption moves from pilot to production, organisations that sized their memory allocation for training are finding that serving is where the constraint actually bites.
The practical implication is that memory demand from AI will continue to outpace supply additions through the forecast period, even under optimistic fabrication assumptions. CTOs should treat any planning model that assumes demand growth slows as a scenario to stress-test, not a base case to rely on.
Cloud vs. On-Premise: Remodelling the Cost Assumptions
The conventional argument for cloud AI infrastructure rests partly on the assumption that hyperscaler purchasing power insulates enterprise customers from component-level scarcity. That argument is weakening. Hyperscalers are prioritising their own AI product buildouts, and reserved capacity for third-party workloads is becoming harder to secure at predictable pricing.
On-premise procurement is not automatically the answer. Organisations that commit to on-premise GPU and memory infrastructure without multi-year supply agreements are exposed to the same spot-market volatility, with less flexibility to shift workloads when pricing spikes. The defensible position is a mixed model, with long-term reserved cloud capacity for baseline inference load and on-premise investment concentrated in workloads where data residency or latency requirements justify the capital commitment.
What changes in a memory-constrained environment is the cost modelling methodology. Reserved cloud pricing needs to be evaluated against a realistic view of what spot and on-demand memory-intensive compute will cost in 2027, not what it costs today. The gap between those two figures is where planning errors accumulate.
Architectural Decisions That Reduce Memory Exposure
The most durable response to memory scarcity is architectural, not procurement-led. Organisations that design AI systems to minimise memory footprint per unit of useful output will be structurally better positioned than those that rely on throwing more capacity at the problem.
Inference Efficiency
Quantisation, speculative decoding, and KV cache eviction strategies can materially reduce the memory bandwidth and capacity required for a given inference throughput target. These are not experimental techniques. They are production-grade approaches that leading inference teams are already deploying to manage cost at scale. The engineering investment required to implement them is real, but it is bounded and predictable in a way that memory procurement costs are not.
Model Selection and Workload Routing
Routing workloads to the smallest model capable of meeting quality requirements is a memory efficiency strategy as much as a cost strategy. Organisations running a single large model for all query types are leaving significant efficiency headroom unused. A tiered routing architecture, where lightweight models handle high-volume, lower-complexity queries and larger models are reserved for tasks that genuinely require them, reduces aggregate memory demand without degrading output quality on the queries that matter.
Building Memory Scarcity Into the Planning Cycle
The structural change that this environment requires is treating memory capacity as a constrained resource in the planning process, with the same discipline applied to headcount or capital expenditure. That means modelling memory requirements explicitly for each AI workload on the roadmap, not deriving them as a residual from GPU provisioning decisions.
It also means building procurement lead times into the project timeline from the start. Memory-intensive hardware acquired through standard procurement channels in a tight market will arrive late, cost more than modelled, or both. Organisations that establish supply relationships and framework agreements before they have an urgent need will have more options than those that enter the market reactively.
The planning horizon that matters here is 2028, not the next budget cycle. CTOs who anchor their infrastructure roadmaps to that horizon, with explicit memory scarcity assumptions built in, will make different and more defensible decisions on cloud commitments, architectural investment, and workload prioritisation than those who treat the current environment as a temporary disruption to be managed quarter by quarter.
Where Vector Labs Fits
We help engineering organisations model infrastructure constraints probabilistically, so that procurement and capacity decisions reflect realistic supply scenarios rather than point estimates. In our memory-wall analysis, we examined how HBM scarcity and KV cache constraints translate directly into inference cost exposure for production AI systems. If you are building a multi-year AI infrastructure plan and want to stress-test the memory assumptions embedded in it, contact us at vector-labs.ai/contacts.
FAQs
Reserved capacity agreements for memory-intensive compute should be evaluated against a 2027 spot-market pricing scenario, not current rates. The gap between today's reserved pricing and projected on-demand costs in a continued tight market is significant, and organisations that defer commitment decisions are likely to find that favourable reserved terms become harder to negotiate as hyperscalers prioritise their own AI workloads. Locking in multi-year reserved capacity for baseline inference load now, while retaining flexibility for burst workloads, is a more defensible position than treating cloud as infinitely elastic.
Prioritisation should be driven by the ratio of business value generated per unit of memory consumed, not by which workloads are most visible internally. High-volume inference workloads with clear revenue or cost-reduction attribution should take precedence over experimental or low-utilisation training runs. Workload routing architectures that direct queries to the smallest capable model also reduce aggregate memory demand, which effectively expands the productive capacity of whatever memory you have provisioned.
Not automatically. On-premise investment without multi-year supply agreements exposes you to the same spot-market volatility as cloud, with less flexibility to shift workloads when conditions change. The stronger hedge is a mixed model: long-term reserved cloud capacity for predictable baseline load, and on-premise infrastructure concentrated in workloads where data residency, latency, or regulatory requirements justify the capital commitment. The key discipline is not choosing one over the other, but ensuring that the cost model for each reflects realistic memory pricing assumptions through 2028.
Quantisation and KV cache management are the highest-leverage interventions for most production inference systems. Quantisation reduces the memory required to hold model weights in active memory, while KV cache eviction and compression strategies reduce the memory consumed per concurrent session in long-context workloads. Both are production-grade techniques with well-understood trade-offs. Speculative decoding can also improve throughput per unit of memory bandwidth, which effectively increases the useful capacity of a fixed memory allocation. The engineering investment to implement these is bounded and predictable, unlike the cost of simply provisioning more memory.
The planning horizon that matters is 2028, based on fabrication ramp timelines for advanced DRAM and HBM capacity. New fabs announced today will not reach meaningful volume output before that window, and HBM packaging constraints add further delay beyond raw wafer capacity. Any roadmap that assumes material supply relief before 2027 is carrying an unpriced risk. The more defensible approach is to build the roadmap on a base case of continued tightness through 2028, with upside scenarios modelled separately, rather than anchoring to an optimistic normalisation timeline that the physical supply chain does not support.

