
According to reports, Micron is exploring the development of high-endurance NAND flash modules and investigating ways to place NAND flash closer to GPUs. The goal is to support workloads that do not require the extremely high bandwidth and low latency provided by conventional HBM or DRAM.
The architecture, currently referred to as “Near-GPU NAND,” is designed to balance storage density and endurance. It would allow GPUs to access a dedicated pool of NAND flash with lower density than conventional TLC or QLC NAND, while potentially improving I/O performance, bandwidth, and read speeds.
The technology could give GPUs direct access to storage pools with capacities of hundreds of gigabytes, potentially reducing memory capacity constraints when running large language models (LLMs). By supplementing HBM and DRAM with high-capacity NAND, the approach could provide a more scalable memory hierarchy for AI workloads.