Kimi K3 model changes the picture, but not how you thought.
The recent launch of the Kimi K3 model caused a stir and declines in chip stocks, similar to what happened after the launch of DeepSeek.
Due to concerns that more efficient models would reduce the need for extensive computing power.
But as more technical details are released, a slightly different picture emerges than what everyone thinks.
Let's start with the interpretation and move to the simple explanation.
The Kimi K3 model is a massive open-source model with 2.8 trillion parameters, a million-token context window, and a Mixture of Experts architecture with a high level of sparsity.
This means that not all parameters are activated for every request, which improves computational efficiency and allows for lower-cost inference.
But there's an important point many are missing: the simple explanation.
Although the model uses less computing power per request, all parameters still need to be stored in memory, occupying about 1.4 terabytes of memory.
Therefore, to operate the model at a large scale, infrastructure with enormous memory capacity and extremely fast memory is required.
Additionally, Kimi K3 is specifically designed for coding tasks and AI agents, tasks that make repeated calls to the model over time and require prolonged inference capabilities.
These workloads make memory bandwidth and HBM capacity a central part of performance.
Therefore, while the demand for computing power per token may improve, the demand for high-quality memory and GPU systems with large amounts of HBM may actually continue to grow.
The reality on the ground also reinforces this claim:
Within hours of its launch, demand for the model exceeded its available capacity, and the company behind Kimi K3 was forced to temporarily halt new subscriptions and allocate existing computing power to paying users only.
On July 27, the company is expected to release the model's weights as open source, meaning any organization will be able to run Kimi K3 itself.
But to do so on a commercial scale will still require significant investment in AI servers and memory-intensive hardware.
Meanwhile, Alibaba also introduced a new model with 2.4 trillion parameters, indicating that competition among Chinese AI companies is only intensifying.
The broader implication is that competition among models might lead to a decrease in AI service prices.
But it could also lead to another wave of investments in infrastructure, servers, GPUs, and memory to run these models for millions of users.
Ultimately,
The debate no longer boils down to how much computing power each model needs, but rather how much memory is required to store and run it at scale.
In the case of Kimi K3, it appears that the advantage in computational efficiency does not eliminate the need for advanced hardware (which caused market concern), but rather shifts the focus of demand towards memory-intensive systems.