- Current Weight Access: Moonshot AI’s Kimi K3 weights are now downloadable. Support from Alibaba Cloud and Huawei’s Ascend AI accelerator ecosystem remains possible and unconfirmed.
- Deployment Work: The SGLang model-serving framework and AI infrastructure provider Baseten demonstrate launch-day compatibility, including eight Nvidia GB300 GPUs per Baseten instance.
- Platform Scope: Alibaba Cloud and Huawei have not established regions, software versions, hardware targets, access terms, or production availability.
- Open-Weight Market: Kimi K3 enters an active market alongside Z.ai’s GLM-5.2 and Thinking Machines Lab’s Inkling open-weight models.
Moonshot AI has released weights for its Kimi K3 AI model after giving early access to AI infrastructure provider Baseten. Moonshot had already made its Kimi K3 large language model available through Kimi.com, Kimi Work, Kimi Code, and its API. Baseten made a Kimi K3 available via API.
Launch-day, or day-zero, compatibility equips serving tools and infrastructure for a model’s release. Kimi K3’s scale makes that work demanding: its checkpoint contains 2.8 trillion total parameters, activates 104 billion parameters, and accepts up to 1,048,576 tokens of input at once. Its mixture-of-experts design activates only part of the model for each token, while the context window measures the text and other input available to a single request.
Total size drives storage needs, while active parameters affect the computing work required for each token. Baseten’s dated API launch and the SGLang model-serving framework provide concrete examples of immediate compatibility, but neither verifies an equivalent Alibaba Cloud or Huawei Ascend managed service. A serving engine can execute the weights, while a production cloud offering must also define accelerator validation, scheduling, access controls, regional capacity, and customer availability.
What Launch-Day Compatibility Requires
Alibaba Cloud and Huawei’s Ascend artificial intelligence accelerator ecosystem also provide day-zero Kimi K3 support. Customers still lack details identifying a managed service, optimized software, validated accelerator configuration, or another deployment path. They also need timing, regions, software versions, hardware targets, supported checkpoints, access conditions, and production availability before treating the support as a deployable service.
Moonshot AI lists vLLM, SGLang, and TokenSpeed as supported inference engines. SGLang added launch-day inference support, while the Miles reinforcement-learning framework added launch-day support for post-training work.
Inference engines manage model execution and incoming requests. Miles addresses a different stage in which developers refine behavior after initial training. Both paths still require accelerator kernels, memory management, scheduling, and an application-facing interface.
Before launch, Baseten used early access to coordinate with the vLLM and SGLang teams. The company built a hosted endpoint with vision input and the full context capacity. Its implementation exposes the deployment burden: Kimi K3’s native weights occupy more than 1.4 terabytes, and one instance uses eight Nvidia GB300 graphics processors.
One provider’s configuration cannot define a universal minimum. Its storage and accelerator requirements demonstrate why launch-day access depends on preparation before public weights arrive. Kimi K3 demand nearly reached available compute, according to Moonshot AI, underscoring why platform support needs capacity details as well as software compatibility.
Open Weights Face a Crowded Field
Kimi K3’s initial launch used an open-weight release format. Developers can obtain the learned parameters and run or adapt the model on their own infrastructure. Open weights are not equivalent to fully open-source software: training data, code, and license terms may remain restricted.
And: Whoever serves Kimi K3 still pays for hardware and electricity.
AI researcher and Interconnects editor Nathan Lambert assesses the gap between leading open and closed models, or between leading American and Chinese models, at roughly three to five months. He identifies Z.ai’s GLM-5.2 and Thinking Machines Lab’s Inkling as relevant alternatives. Thinking Machines Lab released Inkling’s full weights on July 15.
Inkling is a 975-billion-parameter model with 41 billion active parameters, while Kimi K3 is larger and activates more parameters. Inkling also requires a multi-accelerator deployment. Its hardware profile differs from Kimi K3’s without establishing a quality ranking.


