Instructions to use GrEarl/Kimi-K3-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use GrEarl/Kimi-K3-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf GrEarl/Kimi-K3-GGUF:Q2_K # Run inference directly in the terminal: llama cli -hf GrEarl/Kimi-K3-GGUF:Q2_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf GrEarl/Kimi-K3-GGUF:Q2_K # Run inference directly in the terminal: llama cli -hf GrEarl/Kimi-K3-GGUF:Q2_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf GrEarl/Kimi-K3-GGUF:Q2_K # Run inference directly in the terminal: ./llama-cli -hf GrEarl/Kimi-K3-GGUF:Q2_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf GrEarl/Kimi-K3-GGUF:Q2_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf GrEarl/Kimi-K3-GGUF:Q2_K
Use Docker
docker model run hf.co/GrEarl/Kimi-K3-GGUF:Q2_K
- LM Studio
- Jan
- vLLM
How to use GrEarl/Kimi-K3-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GrEarl/Kimi-K3-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GrEarl/Kimi-K3-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/GrEarl/Kimi-K3-GGUF:Q2_K
- Ollama
How to use GrEarl/Kimi-K3-GGUF with Ollama:
ollama run hf.co/GrEarl/Kimi-K3-GGUF:Q2_K
- Unsloth Studio
How to use GrEarl/Kimi-K3-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for GrEarl/Kimi-K3-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for GrEarl/Kimi-K3-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for GrEarl/Kimi-K3-GGUF to start chatting
- Pi
How to use GrEarl/Kimi-K3-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GrEarl/Kimi-K3-GGUF:Q2_K
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "GrEarl/Kimi-K3-GGUF:Q2_K" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use GrEarl/Kimi-K3-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GrEarl/Kimi-K3-GGUF:Q2_K
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default GrEarl/Kimi-K3-GGUF:Q2_K
Run Hermes
hermes
- OpenClaw new
How to use GrEarl/Kimi-K3-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf GrEarl/Kimi-K3-GGUF:Q2_K
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "GrEarl/Kimi-K3-GGUF:Q2_K" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use GrEarl/Kimi-K3-GGUF with Docker Model Runner:
docker model run hf.co/GrEarl/Kimi-K3-GGUF:Q2_K
- Lemonade
How to use GrEarl/Kimi-K3-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull GrEarl/Kimi-K3-GGUF:Q2_K
Run and chat with the model
lemonade run user.Kimi-K3-GGUF-Q2_K
List all available models
lemonade list
- Atomic Chat
Author Note
The author does not own enough hardware to run this 864.81 GiB model. Full-model runtime results below were contributed by independent users and have not been reproduced by the author. They are reported with their environment and limitations rather than presented as author-run benchmarks.
Kimi-K3 GGUF — Q2_K experts / Q4_K dense (2.673 bpw)
864.81 GiB in 94 parts. Text model only. The current v3 payload was
promoted on 2026-08-02 UTC. It uses the same
llama.cpp PR #26185
(pwilkin/kimi-k3-text) tensor contract as the earlier builds.
Source: moonshotai/Kimi-K3 at revision
9f62e4e9fffbd0a83ddd60e1c209d828994b3569 (pinned for every download).
| Parts | 94 (Kimi-K3-Q2_K-000NN-of-00094.gguf) |
| Size | 864.81 GiB (source 1,453.8 GiB → 0.595x) |
| Tensors | 2,573 (matches the contract exactly) |
| Effective rate | 2.673 bpw |
| Architecture | kimi-k3 |
general.file_type |
10 (MOSTLY_Q2_K) |
| Chat template | embedded, validated byte-exact against K3's own builder |
| Quantization release | v3 |
| Sub-block scale search | 15-step ladder + 18 polished multi-start restarts, unweighted squared error |
| Previous release | v2; archived under v2/ at revision a62667d01793 |
| imatrix | no |
Quantization layout
| ggml type | tensors | size | bpw | what |
|---|---|---|---|---|
Q2_K |
276 | 832.04 GiB | 2.6250 | routed expert stacks (ffn_{gate,up,down}_exps) |
Q4_K |
1,067 | 28.58 GiB | 4.5000 | dense 2-D weights, token_embd |
F32 |
1,112 | 2.25 GiB | 32 | norms, router, ssm_a, ssm_conv1d, fused AttnRes scores |
Q8_0 |
1 | 1.16 GiB | 8.5 | output.weight |
F16 |
117 | 0.76 GiB | 16 | attn_k_b / attn_v_b (row length not a multiple of 256) |
The experts hold 2.72 T of the 2.78 T parameters, so they set the overall rate.
Dense weights are kept at Q4_K rather than Q2_K: they are 3.3% of the bytes, and
spending 4.5 bpw there is cheap insurance for the layers that every token passes
through. output.weight is promoted further, to Q8_0, because its error lands
directly on the logits with nothing downstream to average it out. The Q8_0 tensor
is 1.16 GiB and adds 0.54 GiB versus keeping it at Q4_K, or 0.06% of the
artifact; upstream llama.cpp promotes the output head for low-bit file types as
well.
Chat template
K3's tokenizer_config.json ships no chat template — conversation assembly
lives in encoding_k3.py as code, so nothing could simply be copied across. The
template embedded here is Xenova's Jinja
port, patched for llama.cpp's
Minja engine and validated against K3's own build_chat_segments():
28 of 28 cases byte-exact. The cases cover normal and multi-turn messages,
thinking_effort, tool declarations, assistant tool calls, tool results, and a
nested response_schema.
The exact template extracted back from this GGUF was also rendered by llama.cpp's
Minja implementation (commit 91f8c9c5): a tools/tool-result case was 1,218 bytes
and a nested-schema case was 621 bytes, both byte-identical to K3's renderer.
Images and batched conversations remain untested.
Provenance and independent conformance. The file in this artifact began with
Xenova's port and was expanded
and patched locally; it is not claimed to be byte-identical to the separate
upstream implementation in Moonshot PR #66.
PR #66 and ChatLint's K3 findings
provide independent, oracle-backed conformance work. The Minja
namespace(items=[]) collision reproduced here was accepted and fixed upstream
in ChatLint commit ecb1a727
and on the PR branch. Against this artifact's 15,053-byte file, ChatLint pinned to
that commit reports 294/294 checks, 0 errors, 0 warnings. ChatLint renders with
transformers-style Jinja2 rather than Minja and checks structural properties, so
it supplements rather than replaces the 28/28 oracle comparison and direct Minja
runs above.
One limitation is shared with PR #66: sandboxed Jinja cannot parse a JSON string
inside tool_calls[].function.arguments. This template emits a valid
<|open|>json type="object" fallback for a non-empty string; callers that require
the reference renderer's per-argument XTML must parse the string into an object
before applying the template.
Measured quantization error
MXFP4 source dequantized to F32, requantized to Q2_K, then dequantized again and compared against the source. 92 expert tensors sampled, one per routed layer:
| metric | value |
|---|---|
| relative RMSE, mean | 0.263402 |
| cosine similarity, mean | 0.964694 |
The superseded v2 build measured 0.276874 relative RMSE / 0.961095 cosine, and the original v1 min/max build measured 0.3309 / 0.951450, on the same 92-tensor comparison. Those are historical baselines, not measurements of the current files.
For reference, the same measurement on the IQ1_S sibling current top32/refit2 payload is 0.455936 / 0.890050 at 528.0293 GiB. This repo is 1.638x larger by file size and noticeably closer to the MXFP4 source in this weight-space measurement. Historical IQ1_S figures of 0.4626 / 0.886692 for the top8/refit1 payload, and 0.5375 / 0.848230 for the earlier ternary-space snap, are superseded and do not describe its current files.
The v3 scale search runs in a fused CuPy kernel and keeps each 16-value sub-block
in registers while evaluating the 15-step candidate ladder, polishing its winner,
and testing 18 independent restarts. Assignment and least-squares re-solving are
alternated three times per candidate; the lowest unweighted squared error wins.
Everything downstream remains standard block_q2_K: 84 bytes per 256 elements,
or 2.625 bpw.
The Q2_K bytes are fed to gguf-py's own dequantizer and compared against the local dequantizer: max absolute difference 0.0, 0 mismatching elements, on synthetic data and a real K3 expert tensor. Before the full run, the fused search was also checked against a scalar implementation on 16 real K3 expert tensors spanning layers 1–91 and experts 0–895; both produced the same measured relative RMSE, with no sampled tensor regressing against v2.
v1 → v2 → v3
One expert tensor was measured in every one of the 92 rebuilt expert-bearing parts on identical source inputs:
| release | sub-block search | relative RMSE, mean | cosine, mean |
|---|---|---|---|
| v1 | direct min/max | 0.3309 | 0.951450 |
| v2 | 6-candidate unweighted SSE (nstep=5) |
0.276874 | 0.961095 |
| v3 (current) | 15-step ladder + polished 18-restart multi-start | 0.263402 | 0.964694 |
Directly comparing v2 and v3:
| metric | v2 | v3 |
|---|---|---|
| relative RMSE, range over 92 parts | 0.275965–0.277595 | 0.263099–0.263567 |
| cosine, range over 92 parts | 0.960898–0.961344 | 0.964649–0.964776 |
- Relative RMSE falls by 4.87%, equivalent to removing 9.50% of squared reconstruction error.
- Cosine similarity to the source weights rises by 0.360 percentage points.
- All 92/92 measured parts improved; 0 regressed. Per-part RMSE reduction is 4.66%–5.06%.
- From v1 to v3, relative RMSE is down 20.4% at the same 2.625 bpw expert payload size.
- File size, tensor count, tensor names, ggml types, and runtime-relevant metadata are unchanged. Only the routed-expert Q2_K payload differs; parts 1 and 94 contain no routed experts and are byte-identical to v2.
The expert-free part 1 was copied server-side during promotion, so its custom
provenance KV kimi-k3.conversion.q2k_subblock_search still says
unweighted-sse nstep=5. That value accurately identifies v2 but is stale for
the current expert payload; the authoritative v3 setting is the 15-step ladder
plus 18 polished restarts documented here. This does not affect loading or
inference.
These are weight-space reconstruction measurements. They show that v3 sits closer to the MXFP4 source, but they do not quantify generation quality. No perplexity, standardized benchmark, or source-model output-equivalence comparison exists for v1, v2, or v3.
The v3 run used eight NVIDIA L4 workers and rebuilt the 92 expert-bearing parts in 36.5 minutes for a recorded Modal cost of $2.635. The fused search itself was verified before publication; every emitted part also passed the no-regression gate and structural checks for its split number, tensor count, tensor types, and maximum tensor-name length.
Previous builds
The v2 files were archived under v2/ during promotion. They remain reachable
in the Hugging Face commit history by pinning revision a62667d01793 and
using the v2/ path, even after that folder is removed from the current file
listing. The original v1 min/max build remains reachable at revision
d76240360965, where it was stored at the repository root.
Previously downloaded v1 or v2 files remain loadable. The measured difference is weight-space error only, with no downstream benchmark comparison, so re-downloading 864 GiB is a quality-versus-bandwidth decision rather than a compatibility fix.
Runtime status: v2 full load reported; v3 structure verified
kimi-k3 is not merged into released llama.cpp; support exists on
PR #26185
(pwilkin/llama.cpp, branch kimi-k3-text).
A detailed independent report in
Discussion #4
(2026-07-28 UTC) exercised the 94-part Q2_K root available at that time: v2, not
the current v3 payload. The repository root was at
d8aebda6c891e203284fc4ddef5483f2177ee841 when the report was posted; the
reporter did not provide per-file hashes, so this attribution is based on the
repository timeline rather than a cryptographic pin. v3 preserves every tensor
name, type, shape, split field and runtime-relevant metadata, but it has not been
separately loaded as a full 94-part model.
Reported environment and results:
| Item | Third-party v2 report |
|---|---|
| Hardware | AWS p6-b200.48xlarge: 8× NVIDIA B200 / 1.43 TiB aggregate HBM, sm_100 |
| Runtime | pwilkin/llama.cpp kimi-k3-text at 06eec9f5 |
| Load | All 94 parts loaded cleanly, fully VRAM-resident; 869 GiB resident, 97–116 GB per GPU |
| Startup | About 110 seconds from launch to serving from local NVMe |
| Decode | 17.5–17.8 tokens/s, reported stable across tasks |
| Prompt processing | Up to about 330 tokens/s with -ub 4096 |
| Generation checks | At temperature 0, the reporter observed coherent factual answers, instruction following, and correct handling of a simple trap question |
| Tools | The model emitted well-formed tool calls with correctly typed arguments |
| Context | A 262K context allocation ran with little VRAM change and unchanged decode speed; needle retrieval was tested at 26K tokens, not 262K |
| Input rendering | The reporter confirmed that the 28/28-tested prompt template rendered correctly in practice |
These are third-party observations of v2, not author-run measurements and not a v3 benchmark. They establish full load, serving, generation, raw tool-call emission, and a 26K retrieval check for v2 on the named setup. They do not establish v3 generation quality, perplexity, benchmark scores, source-model equivalence, or 262K retrieval quality. The statement that no quantization damage was visible is a qualitative user observation, not a formal quality result.
Known runtime limitations from the same report
- CPU-only currently crashes while loading K3 at
GGML_ASSERT(*cur_backend_id != -1)inggml_backend_sched_split_graph. At present use a GPU or GPU+CPU hybrid configuration; the traditional all-CPU large-RAM/mmap route is not validated. - The tested branch lacked a K3 output parser. Without the reporter's
parser PR,
/v1/chat/completionscan return emptycontent, place output inreasoning_content, and leak raw XTML markers. This is an output-parsing issue, not evidence that the embedded input template is wrong. special_eos_id is not in special_eog_idswas reported as a benign warning; raw/completionstopped withstop_type: eos. The metadata has not been changed solely to silence that warning.- Released/stock llama.cpp, CPU-only execution, Metal, Vulkan, and other hardware configurations remain unverified for this artifact.
What has been verified, and by whom:
| Level | Status |
|---|---|
| GGUF structure, split metadata, tensor names/types, hparams, vocab | Verified by the author on the published files via HTTP Range reads |
| Numeric agreement of custom quant bytes with gguf-py's dequantizer | Verified by the author, exact |
Chat input template vs K3 build_chat_segments() |
Verified by the author, 28/28 byte-exact; tools/schema rendered directly with Minja |
| Full 94-part load, serving, generation, tool emission on 8×B200 | Reported for v2 by an independent user in Discussion #4; v3 not separately tested |
| 262K context allocation / 26K needle retrieval | Reported for v2 by the same user; retrieval was not tested at 262K |
| Perplexity, standardized benchmark, source-model output equivalence | Not measured |
Template compatibility fix
The first template published here used Jinja's namespace({...}) dict-literal
form. Minja accepts only namespace(key=value, ...). Direct Minja execution then
found two more tool/schema-path issues: tojson(sort_keys=true) is not implemented,
and a namespace field named items collides with the object method. The current
15,053-byte template uses recursive dictsort JSON rendering and a non-conflicting
field name. It is embedded in part 1 and also published as chat_template.jinja.
Thanks to the user who caught the original incompatibility.
What was fixed relative to the previous 96-part upload
The earlier Q2_K release in this repo (96 parts, 938.59 GiB) could not have been loaded even with K3 support present. All four defects required regenerating every part, so this upload replaces it entirely.
- Tensor names were raw Hugging Face names. 668 names across 92 of the 96
parts exceeded
GGML_MAX_NAME(64). Now every name follows the PR contract; longest is 29 characters. - Part 1 had no model metadata. No hparams, no vocabulary. Now part 1 carries
66 KV entries: 25 required hparams plus the full vocabulary (163,840 tokens /
163,328 merges,
gpt2/ prekimi-k2, chkhsh verified). split.tensors.countwas UINT16 and per-part.llama-model-loader.cppreads it as a required key and compares it againstweights_map.size(), throwingcorrupted modelon mismatch. Now INT32 / 2573 in all 94 parts, equal to the actual sum of per-part tensor counts.- Vision tensors were mixed into the text model. HF shards 95 and 96 hold 168
vision_tower/mm_projectortensors.done_getting_tensors()is called withpartial=false, so a single extra tensor triggerswrong number of tensors. Those two shards are excluded, which is why there are 94 parts and not 96.
Conversion details
Recorded in the file itself under kimi-k3.conversion.*:
a_log_transform=-exp(A_log[:n_head])— K3 storesA_logas 128 elements but only the firstnum_attention_heads= 96 are usedkv_b_split=k_b(transposed)+v_b—kv_b_projsplit intoattn_k_b/attn_v_battn_res_fused=res_norm*res_proj[0] in float32— the AttnRes norm/proj pair is fused into one score vector, matching the reference_apply_attn_res()expert_stack_dim= 0 — 896 experts stacked on dim 0source_quant=compressed-tensors/mxfp4-pack-quantizeddefaults_used=rope_theta— the only value not present inconfig.json, taken from theconfiguration_kimi_k3.pyclass default (10000.0)
Usage on the PR branch
Pass the first part; llama.cpp follows the split metadata to the rest.
llama-cli -m Kimi-K3-Q2_K-00001-of-00094.gguf -p "..."
All 94 parts must be present in the same directory.
- Downloads last month
- 4,286
2-bit
Model tree for GrEarl/Kimi-K3-GGUF
Base model
moonshotai/Kimi-K3