• Skip to main content
  • Skip to secondary menu
  • Skip to footer

Technologies.org

Technology Trends: Follow the Money

  • Technology Events 2026-2027
  • Sponsored Post
  • Technology Markets
  • About
    • GDPR
  • Contact

Why DRAM and HBM Demand Grows as AI Matures, and Where the Cycle Still Bites

August 15, 2026 By admin Leave a Comment

The standard assumption about maturing technology is that hardware demand flattens. Software gets more efficient, silicon gets denser, and the same work needs fewer parts. That has been broadly true for compute. It has not been true for memory, and the reasons are structural rather than cyclical.

The short version: as AI moves from training to inference, the binding constraint moves from arithmetic to data movement. Memory sits on the wrong side of that shift, which is to say the profitable side.

Inference Is a Bandwidth Problem

Generating a token requires streaming the active model weights out of memory and into the accelerator. The arithmetic involved is trivial by modern standards. The data movement is the entire cost. A GPU waiting on memory is a GPU doing nothing, which is why accelerator vendors keep adding stacks rather than adding cores.

Training is episodic and capex-lumpy. It happens in bursts, tied to model release schedules and cluster build-outs. Inference is continuous and scales with adoption. TrendForce expects inference to overtake training as the primary driver of AI server demand before the end of the decade, and every forecast that moves in that direction moves demand toward bandwidth and away from raw compute.

The State Scales With Users, Not With Models

This is the part most demand models understate. Every active session holds a key-value cache, and that cache grows linearly with context length and with the number of concurrent users. It has nothing to do with how large the model is.

Three things are pushing it hard right now. Context windows keep expanding. Reasoning models generate an order of magnitude more tokens per query than single-shot answers did, and every one of those tokens extends the cache before the user sees a word. Agentic workloads hold sessions open for minutes or hours instead of seconds.

So even in a world where model architectures froze tomorrow, per-user memory consumption would keep climbing on adoption alone. Industry estimates put real-time memory demand across the major inference platforms at roughly 750 petabytes, and closer to 1.5 exabytes once you count the redundancy and headroom any production deployment actually requires.

Efficiency Moves Demand, It Does Not Remove It

Each efficiency technique that gets cited as a threat to memory demand turns out, on inspection, to relocate it.

Mixture-of-experts architectures explicitly trade compute for capacity. You hold every expert resident in memory and activate a small fraction per token. Sparse models are more memory-hungry per unit of compute, not less. Quantization cuts the cost per token, which raises token volume, which is the oldest pattern in the industry. Distillation produces smaller models that get deployed in far more places.

The classic Jevons shape applies. Cheaper inference means more inference.

The Wafer Math Amplifies Everything

Bit shipments understate what AI is doing to supply, because high-speed memory is far more expensive to manufacture per bit. A gigabyte of HBM consumes roughly four times the fab capacity of standard DRAM once you account for die area, through-silicon vias, and stacking yield. GDDR7 runs around 1.7 times.

Total DRAM wafer starts are growing something like six to eight percent annually. Every wafer allocated to HBM is a wafer not producing conventional DDR5. That crowding-out effect is why commodity DRAM pricing has moved with the AI cycle despite having almost nothing to do with AI workloads, and why memory has started behaving like an allocation market rather than a spot commodity market.

Where the Thesis Can Still Cost You

Structural demand and a non-cyclical stock are different objects, and conflating them is how people get hurt in this sector.

The supply response is late, not absent. New fabs take two to three years from commitment to meaningful output. The historical failure mode has never been that demand disappeared. It is that peak capacity arrived into a demand pause. Korean cluster announcements and Micron’s Japanese expansion will not add bits before 2027 or 2028, which supports the near term and complicates the one after it.

CXMT is currently more useful as a negotiating lever for buyers than as a genuine supply threat. It lags badly in HBM, it lacks EUV, and its mobile mix skews to older LPDDR generations. That could change slowly. It is unlikely to change suddenly.

The real technical risk to the thesis is architectural. Attention variants that carry constant-size state, or aggressive key-value compression, would cut per-token memory intensity in a way no process node ever could. Low probability on a two-year view. Not zero on a five-year one.

For anyone watching the cycle rather than the secular story, the 2027 HBM4 contract negotiations running through the second half of this year are the first real price discovery since suppliers sold out. Firm pricing validates the extended-supercycle case. Any concession pattern shows up two to three quarters before it appears in reported results.

The level rises structurally. The cycle still exists. Those are separate trades.

Filed Under: News

Reader Interactions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Footer

Recent Posts

  • Why DRAM and HBM Demand Grows as AI Matures, and Where the Cycle Still Bites
  • The AI Boom Is Broadening: Intel’s $100 Billion Book, CoreWeave’s 1.5 Gigawatts, and Gemini’s Billionth User
  • Autodesk Opens Fusion to AI Agents as SendCutSend Banks $110 Million: The Design-to-Part Loop Is Now Machine-Readable
  • Cloudflare Open-Sources Cloudflare OS: The Agent Workspace Is Free, the Network Underneath Is Not
  • Marvell (MRVL) Turns Celestial AI Into Product, and the $5.5 Billion Earnout Clock Is Now Running
  • Samsung Unveils zHBM and 400-Layer V10 BV-NAND at FMS 2026, and Wafer Bonding Is the Common Thread
  • Kioxia GP1 Wins FMS Best of Show With 10 Million IOPS, Splitting From SanDisk and SK Hynix on HBF
  • The Humanoid Robot Bottleneck Is the Battery: Why Two Kilowatt-Hours Caps the Whole Industry
  • SK hynix HBF Standard Turns NAND Into a Memory Tier, and the Memory Trade Still Has Room to Run
  • The Humanoid Trap: FCC Robot Import Ban Defends the Wrong Form Factor

Media Partners

  • Market Analysis
  • Cybersecurity Market
  • App Coding
US Market Cap at $74 Trillion: Why the Market Has Room to Grow Without Repricing
Buffett Indicator at 230%: Why the Labor Share Makes Market Cap to GDP Unreadable
60-Month Transformer Lead Times Are a Bigger AI Constraint Than the Copper Deficit
America Mines the World’s Semiconductor Quartz and Has No Export Controls on It
China’s Equipment Export Controls Are the Real Threat to America’s 2028 Magnet Timeline
Earnings Recap August 3-7, 2026: AMD, Datadog and SanDisk All Beat and All Fell
July Jobs Report: The 103,000 Revised Away Matters More Than the 23,000 Lost
Big Tech Capex Reaches $1.1 Trillion Since 2023, With $745 Billion Planned for 2026
Amphenol’s Record Quarter Shows Where AI Capex Actually Lands
Paper Raises $34 Million and Figma (FIG) Has Already Lost Half Its Value on the Thesis
Oligo Security Raises $60 Million as Runtime Vendors Turn Post-Mythos Into a Market Category
ISACA Europe Conference 2026: AI Governance and Cyber Resilience in Munich, 7-9 October
Bitdefender Adds EU-Only MDR to Its Sovereign Acceleration Program, Turning Data Sovereignty Into a Product SKU
Lattice Semiconductor Closes $1.65 Billion AMI Acquisition, Merging Server Firmware With Root-of-Trust Silicon
NVD Hits 45,207 Flaws in 2026 as Microsoft Prices AI Vulnerability Discovery at Half the Market
Way Security Raises $20M Seed From Insight Partners and Glilot for AI-Driven Identity Deployment
Jensen Huang Is Right About Open Models and Wrong About Cybersecurity
Glow Emerges From Stealth With $180 Million Series A At $1.2 Billion Valuation
Cisco Releases Antares-350M and Antares-1B Open-Weight AI Models for Vulnerability Detection
OpenAI Models Breached Hugging Face Infrastructure While Cheating on Cybersecurity Benchmark
Cloudflare Kitesurf: An Agent-First Browser That Uses 3-7x Less Memory Than Chromium
Vibe Coding Works Until You Have to Read the Code
Asynchronous Programming in Python: How the Event Loop, Event Queue, and Thread Pool Fit Together
PixVerse Closes Series C Extension at $439 Million and Pivots From AI Video Into Games
DigitalOcean Launches AI-Native Cloud at Deploy 2026
Verdent Updates AI Platform to Function as a Full Engineering Team for Solo Builders
The Side Project App Is Not Dead. The Side Project App Business Is.
The App Monetization Landscape Has Changed and Most Teams Have Not Caught Up
Building Offline-First Mobile Apps Is Harder Than It Looks and Worth It
State Management in React Native Has Too Many Options and One Right Answer

Media Partners

  • Market Research Media
  • Technology Conferences
  • API Coding
Weekly Network Analytics, July 19 to July 25, 2026: Visits Up 14%
Adobe (ADBE) and Figma (FIG) Have Each Lost Roughly Half Their Value to a Competitor Set Worth $34 Million
Getty Images Kills the $3.7 Billion Shutterstock Merger Rather Than Sell the Editorial Business the UK Demanded
Fox’s $22B Roku Deal: 4.6x Sales, Paid in 1.5x Stock
Tuesday Open: AI Earnings Engine Holds the Line as Iran Overhang Fades to Noise
China’s U.S. Treasury Holdings: The Great Repositioning (2021–2025)
Infographic: Why the 2025 CIPA Data Proves the APS-C Renaissance is Real
How WiFi Changed Media
Canva Acquires Simtheory and Ortto to Build End-to-End Work Platform
Netflix Price Hikes, The Economics of Dominance in a Saturated Streaming Market
Q4 2026 Semiconductor and Memory Conferences: Dates, Locations, Who Presents
FMS 2026 in Santa Clara: Kioxia, Samsung, SanDisk and SK Hynix Offer Four Incompatible Fixes for the AI Memory Wall
San Francisco AI Summit 2026: Korea-US AI and Semiconductor Summit, July 24, San Francisco, California
SIGGRAPH 2026 in Los Angeles: NVIDIA’s Physical AI Day, a First Games Summit, and the Bolt Graphics Zeus Bet
Inside AMD Advancing AI 2026: Lisa Su Puts Helios on Stage as OpenAI, Meta, Anthropic and Cerebras Line Up Behind It
Remaining 2026 Tech Conferences: Black Hat, Dreamforce, Web Summit Lisbon and AWS re:Invent
2026 Esri User Conference — July 13–17, San Diego
HubSpot UNBOUND 2026: Analyst Day Set for September 17 in Boston
The Signal for the Event-Tech Sector
The 10 Most Significant Tech Events and Earnings to Watch This Summer
Every Accident in Your API Becomes a Contract
Why Private Domain Data Is the Real Key to AI That Actually Works
Orkes Raises $60M to Bring Production-Grade AI Orchestration to Enterprise Developers
Form.io Launches MCP Server and Agentic Coding Toolset for Governed Enterprise AI Development
Appdome Upgrades MobileBOT Defense With Identity-First Mobile API Protection
Five SDK Generators Compared: Speakeasy, Stainless, Fern, APIMatic, and OpenAPI Generator
API Monetization Models That Work and the Ones That Drive Developers Away
gRPC in Production: What the Documentation Doesn't Tell You
Event-Driven Architecture vs Request-Response: Choosing the Right Communication Pattern
The Business Case for Internal APIs That Most Engineering Leaders Ignore

Copyright © 2026 Technologies.org

Media Partners: Market Analysis · Market Research · Referently · Photography