Kimi-K3 Open-Weight Frontier Model | Generated by AI

Home 2026.07

Question: What is moonshotai/Kimi-K3 on Hugging Face?

Answer:

Moonshot AI’s Kimi-K3 is their latest open-weight frontier large language model released on Hugging Face. It is basically Moonshot’s answer to models like GPT-5-class systems and Claude-class models, but with downloadable weights. (Hugging Face)

The important part is not just “2.8 trillion parameters” — the architecture is what makes it interesting. (Hugging Face)

1. Architecture

Kimi K3

2.8T total parameters
        |
        v
Mixture-of-Experts (MoE)
        |
        +-- 896 experts
        |
        +-- each token activates only 16 experts
        |
        v
~104B active parameters/token

So it is not like a dense 2.8T model.

A dense model:

token
 |
all 2.8T parameters run
 |
huge compute

K3:

token
 |
router
 |
choose 16 / 896 experts
 |
only ~104B parameters execute

This is the same scaling idea used by models like Mixtral, DeepSeek MoE, etc. (Hugging Face)


2. Key specs

From the model card: (Hugging Face)

Item Kimi K3
Architecture MoE
Total params 2.8T
Active params 104B
Experts 896
Experts/token 16
Context 1,048,576 tokens
Vision Yes
Layers 93
Attention Kimi Delta Attention + Gated MLA
Weight format MXFP4
Vocabulary 160K

3. Why Kimi Delta Attention (KDA) matters

The normal Transformer attention:

[ Attention(Q,K,V)=softmax(\frac{QK^T}{\sqrt d})V ]

has:

sequence length N

memory:
O(N²)

For million-token context:

1,000,000 tokens

attention matrix:
1e12 entries

Impossible.

Kimi introduces Kimi Delta Attention, trying to make long-context processing cheaper by compressing attention state. (Hugging Face)

Conceptually:

Traditional:

token1 token2 token3 ... token1000000

every token attends every token

KDA:

token stream
    |
compressed memory state
    |
retrieve relevant history

Similar direction to:

but integrated into a Transformer architecture.


4. Why people are excited

The big shift:

Before:

Frontier model
=
OpenAI / Anthropic / Google
=
closed weights

Now:

Frontier-level capability
+
downloadable weights
+
self hosting

K3 joins the trend:

where researchers can inspect, modify, quantize, and deploy locally. (GitHub)


5. Can your RTX 4070 run it?

Short answer: no, not normally.

Your RTX 4070:

VRAM: 12GB

K3 checkpoint:

~1.5TB+

The full model is split into many safetensor shards. (Hugging Face)

But because it is MoE, clever systems can partially load experts:

CPU RAM
 |
model weights
 |
GPU
 |
active experts

Possible experiments:

Community experiments have already shown unusual low-GPU deployments because only selected experts are active. (Reddit)

For practical local inference:

RTX 4070 12GB
    |
    +-- Qwen3-8B
    +-- Qwen3-14B quantized
    +-- Llama 70B via CPU offload (slow)

Kimi K3
    |
    +-- server cluster
    +-- many GPUs

6. Why this matters for AI engineers

For someone learning vLLM / inference stack, K3 is actually a very good architecture study case.

You can learn:

MoE routing

hidden state
      |
      v
router()
      |
      +--> expert 13
      +--> expert 201
      +--> expert 777

Expert parallelism

GPU0:
 expert 0-100

GPU1:
 expert 101-200

GPU2:
 expert 201-300

KV cache pressure

1M context makes:

KV cache
+
attention memory
+
paging

the main bottleneck.

This is exactly the problem vLLM is solving.


My take: Kimi K3 is less interesting because “2.8T” and more because it is a MoE + long-context + agent-oriented architecture. It represents the direction where future frontier models may look like:

huge sparse model
+
cheap active compute
+
persistent memory
+
tool agents
+
local/private deployment

which is very aligned with the inference engineering work you are currently studying.

References:


Generated by AI. Curating and sharing still takes effort. If you find it useful, feel free to donate. WeChat: @lzwjavaWeChat QR · X: @lzwjava · Say hi 👋

Back Donate