The Open-Weight Onslaught: DeepSeek V4, Kimi K3, and the Erosion of the Closed-Source Moat
Event Core
A rapid-fire succession of major releases from DeepSeek, Moonshot (Kimi), Liquid AI, and Mistral signals a paradigm shift where open-weight models are no longer just “catching up” but are setting the pace for architectural innovation and inference efficiency.
- ▶ DeepSeek V4’s Efficiency Play: By leveraging the MXFP4 (Microscaling Formats) precision architecture within a Mixture-of-Experts (MoE) framework, DeepSeek is drastically lowering the barrier for SOTA-level reasoning and massive context handling.
- ▶ The Post-Transformer Era: The arrival of Liquid AI’s non-Transformer foundation models suggests a structural diversification of the GenAI stack, targeting the inherent memory and compute bottlenecks of traditional attention mechanisms.
- ▶ Strategic Convergence: The imminent launch of Kimi K3 and rumors of GLM 5.5 indicate that leading Chinese labs are weaponizing open-weight strategies to challenge the dominance of proprietary Silicon Valley models.
Bagua Insight
At 「Bagua Intelligence」, we view this as the “Commoditization of Intelligence.” The moat surrounding closed-source providers is evaporating as the delta between proprietary APIs and open-weight models shrinks to negligible levels. DeepSeek’s push into MXFP4 is particularly disruptive; it’s a hardware-software co-design move that squeezes maximum “intelligence per watt.” Furthermore, the rise of alternative architectures like Liquid AI proves that the industry is moving beyond the Transformer monoculture. We are shifting from a race of “who has the most GPUs” to “who can do the most with the GPUs they have.”
Actionable Advice
Enterprises should pivot from a “Closed-First” to a “Hybrid-Open” AI strategy. The cost-to-performance ratio of deploying quantized MoE models locally is becoming too significant to ignore. For technical leads, the priority should be mastering the deployment of MXFP4-optimized stacks and exploring non-Transformer alternatives for long-sequence tasks where traditional KV cache costs become prohibitive.