Samsung Electronics is setting its sights on a significant leap in AI performance, aiming to boost the response speed of AI accelerators by a factor of ten. The ambitious goal is spearheaded by the company’s new generation AI memory, zHBM, unveiled recently.
Revolutionizing AI Memory
At the “AI Infrastructure Summit” held in Santa Clara, California, Samsung revealed its roadmap for the future of AI memory. Kim In-dong, a senior executive at Samsung Semiconductor America responsible for storage product planning, shared the company’s vision. “Existing conversational AI systems process 100 tokens per second per user. Our goal is to jump to 1000 tokens per second to support intelligent agents,” he stated.
The key to this advancement lies in Samsung’s zHBM technology. This innovative design involves vertically stacking High Bandwidth Memory (HBM) directly onto AI accelerators. Kim likened this to installing a “dedicated express elevator directly to the lobby from a hotel room.” This architectural shift is expected to deliver performance up to 8 times that of HBM5 and triple its energy efficiency.
Beyond zHBM, Samsung also announced plans to begin providing samples of its three-dimensional storage solution, zNAND-O, starting in 2028. This indicates a long-term commitment to pushing the boundaries of memory technology.
Debate Around CXL Technology
While Samsung focuses on its advanced HBM solutions, the summit also saw discussions surrounding Compute Express Link (CXL) technology. Often considered a supplementary solution to HBM, CXL faced skepticism from some attendees.
CXL is a next-generation memory interconnect standard designed to allow computing devices to connect to more memory, enabling server expansion and memory sharing across different servers. AI accelerators access this memory via the CPU.
Some in the semiconductor industry remain optimistic about CXL, believing its widespread adoption could alleviate issues of high HBM prices and supply shortages. The argument is that servers could maintain similar performance with reduced HBM usage, significantly lowering infrastructure costs.
However, Daniel Morris, a researcher focused on AI accelerator design at OpenAI, expressed reservations. “For the actual running of AI models, I can’t find where CXL would be used. If I had to find one use, it would be for storing infrequently accessed cold data in large AI models. The outside world has expectations for it, but I haven’t seen practical applications specifically for model operation yet,” he commented.
Adding to the debate, Vidya Thiagarajan, Head of Intel’s AI System-on-Chip (SoC) Architecture, highlighted the limitations of CXL. “While consolidating memory with CXL is beneficial, it supplements auxiliary storage and cannot replace HBM. Data transfer between GPUs via CXL is far slower than HBM,” she noted. This suggests that CXL may struggle to displace HBM in performance-critical applications.
Samsung and SK Hynix Poised to Lead
The ongoing advancements in HBM technology, coupled with the limitations of alternative solutions like CXL for core AI processing, suggest that market leaders Samsung Electronics and SK Hynix are well-positioned to maintain their strong momentum in the high-bandwidth memory market. Samsung’s aggressive push with zHBM underscores its commitment to dominating the next era of AI infrastructure.








