Mark RadarMARK RADAR
EN

AI Inference Reshapes HBM Economics as Token Throughput Formula Explains Memory Growth

1 reports · First detected 2026-05-04 · Last active 2026-05-04

Traditional DRAM demand rises and falls with PC sales and economic cycles, but AI inference shifts the benchmark to the number of tokens generated per unit of cost and power consumption. Semiconductor architecture analyst fin said batch size is constrained by HBM capacity, while generation speed is limited by bandwidth. The product of the two therefore becomes central to GPU throughput and memory value.

Blockchain news outlet ABMedia reported on May 4, 2026, that if Nvidia’s token throughput doubles with each GPU generation, the product of HBM capacity and bandwidth must also double. The trend line from the A100 to Rubin Ultra closely tracks that relationship. Supply is concentrated among SK hynix, Samsung and Micron. The original article did not cite procurement spending and cautioned that capacity expansion could still restart the pricing cycle.

All Coverage

1 original reports

The Backstory

The history behind this event

No historical echoes for this signal

Mark Radar|MARK RADAR
All times are in Taipei time (GMT+8)