▶ Full Post Text
The high price burden of Nvidia, which dominates the graphics processing unit (GPU) market, is pushing Japan's artificial intelligence (AI) industry to actively seek alternatives, opening new opportunities for South Korean AI semiconductor companies. In particular, neural processing units (NPUs) specialized for inference rather than training are gaining attention as next-generation alternatives, leveraging their power efficiency and price competitiveness.
According to a report by the *Nikkei* on the 15th, the price of Nvidia AI servers, which can cost hundreds of millions of won per unit (hundreds of thousands of USD), has emerged as the biggest obstacle to AI data center construction plans in Japan. While Nvidia holds roughly 90% of the global market share for data center GPUs, critics argue that the cost and supply chain burden are too great to meet the exploding demand for AI infrastructure.
In response, Japanese companies are actively exploring options beyond simply swapping GPUs for other chips, including improving GPU utilization efficiency through memory optimization and adopting low-power, low-cost semiconductors specialized for AI inference.
The Japanese subsidiary of U.S.-based Penguin Solutions plans to launch its "Memory AI KV Cache Server" in the Japanese market within the year, addressing the chronic problem of memory bottlenecks in GPU-equipped servers. When a server's memory data transfer speed and capacity cannot keep pace with the GPU's computation speed, overall system efficiency drops sharply. This solution stores the KV cache used by large language models (LLMs) in external memory, reducing GPU usage and significantly improving cost efficiency. Penguin Solutions also has a strategic partnership with South Korea's SK Telecom, drawing attention to the potential for AI infrastructure cooperation between South Korea and Japan.
The entry of South Korean AI semiconductor startup Rebellions into the Japanese market is also becoming visible. According to the *Nikkei*, Tomen Devices, Japan's largest semiconductor and electronic components distributor, has begun proof-of-concept testing of servers equipped with Rebellions' NPUs in collaboration with a local AI company. NPUs are semiconductors specialized for inference computations performed during the actual service phase rather than AI model training, and are considered to offer superior power efficiency and price competitiveness compared to GPUs.
Kiyotaka Nakao, president of Tomen Devices, expressed optimism in an interview with the *Nikkei*, stating that "NPUs will become a strong option for building AI infrastructure."
The *Nikkei* analyzed that this expansion of NPU adoption, along with the "post-Nvidia" movement involving Google's internally developed tensor processing units (TPUs), could become a significant variable affecting GPU supply-demand dynamics and market landscape over the medium to long term.
Industry observers expect NPU demand to expand further as the AI industry's center of gravity rapidly shifts from model training to inference. While GPUs maintain their status as the core infrastructure for large-scale AI model training, the inference market is seeing NPUs, company-specific custom AI chips, and memory optimization technologies emerge as new alternatives, likely further diversifying the competitive landscape for AI semiconductors.
This movement in the Japanese market aligns with a global trend to reduce dependence on Nvidia GPUs, and is expected to serve as an important catalyst for expanding overseas market opportunities for South Korean AI semiconductor companies.
[https://finance.biggo.com/news/8f3bf440-cefd-4102-ae90-0de463f36e52](https://finance.biggo.com/news/8f3bf440-cefd-4102-ae90-0de463f36e52)