
Tech • IA • Crypto
China’s latest AI surge shows that compute capacity, not model quality, is becoming the main bottleneck as companies race to deploy trillion-parameter systems at scale.
Moonshot AI launched Kimi K3, a 2.8 trillion-parameter open-weight model focused on coding and agent tasks, quickly drawing global demand. Within 48 hours, usage exceeded forecasts, forcing the company to pause new subscriptions as GPU capacity struggled to keep up. Existing users retained access while expansion plans were accelerated.
Unlike simple chatbots, agent-based systems repeatedly call models to plan, execute, debug, and refine tasks. This dramatically increases compute usage per query. At scale, millions of such workflows push infrastructure limits, making large models significantly more expensive to serve than to train.
K3’s weights are expected to be released, but its size makes independent deployment unrealistic for most users. Running a model of this scale requires vast high-end hardware, meaning cloud access will remain the dominant distribution method despite its open-weight status.
Founded in 2023 by Yang Zhilin, Moonshot has raised over $5.5 billion, including a $2 billion round backed by firms like Meituan and China Mobile, and reportedly reached a $30 billion valuation. The company is preparing for a potential Hong Kong IPO, increasing urgency to secure compute resources and scale reliably.
Analysts note that while Chinese firms are closing the AI performance gap with the U.S., access to advanced chips remains limited due to export restrictions. Delivering models globally at speed is harder than building them, and users have already reported slower response times compared to top Western systems.
AI-related stocks moved sharply following K3’s release. ChinaSoft surged nearly 23% after announcing a partnership with Moonshot, while rivals like Zhipu AI and MiniMax saw double-digit declines. Investor sentiment appears highly sensitive to each major model launch.
DeepSeek V4 is expected to compete closely with frontier models while undercutting them on cost. Pricing could drop to $0.28–$1.74 per million output tokens, far below some competitors charging up to $50. A new peak/off-peak pricing model may push enterprises to shift workloads to cheaper time windows.
Alibaba unveiled Qwen 3.8 Max, a 2.4 trillion-parameter multimodal model using a mixture-of-experts design to reduce compute load per query. It targets enterprise use cases such as coding, data analysis, and automation, and is being integrated into Alibaba’s cloud and agent platforms.
Reports suggest DeepSeek V4, Qwen 3.8, and even Qwen 4.0 could launch within months of each other. Additional models like GLM 5.3 are also in development, indicating an accelerating release cycle where advantages may be short-lived.
The growing use of model distillation allows smaller systems to mimic larger ones at lower cost, but it also introduces concerns around intellectual property and transparency, especially when links to proprietary models are suspected but unconfirmed.
China’s AI race is shifting from building powerful models to sustaining them at scale, where compute access, cost efficiency, and infrastructure readiness are becoming the decisive factors.