Instructions to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0 # Run inference directly in the terminal: llama cli -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0 # Run inference directly in the terminal: llama cli -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Use Docker
docker model run hf.co/CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
- LM Studio
- Jan
- vLLM
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CanadaDiver/KiCad-Gemma4-E2B-v3-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CanadaDiver/KiCad-Gemma4-E2B-v3-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
- Ollama
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with Ollama:
ollama run hf.co/CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
- Unsloth Studio
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for CanadaDiver/KiCad-Gemma4-E2B-v3-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for CanadaDiver/KiCad-Gemma4-E2B-v3-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for CanadaDiver/KiCad-Gemma4-E2B-v3-gguf to start chatting
- Pi
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Run Hermes
hermes
- OpenClaw new
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with Docker Model Runner:
docker model run hf.co/CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
- Lemonade
How to use CanadaDiver/KiCad-Gemma4-E2B-v3-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull CanadaDiver/KiCad-Gemma4-E2B-v3-gguf:Q8_0
Run and chat with the model
lemonade run user.KiCad-Gemma4-E2B-v3-gguf-Q8_0
List all available models
lemonade list
- Atomic Chat
KiCad-Gemma4-E2B-v3 GGUF
Quantized GGUF of CanadaDiver/KiCad-Gemma4-E2B-v3 — a KiCad EDA specialist fine-tuned from google/gemma-4-E2B-it to generate .kicad_sch schematics, write KiCad Python scripts, select components, and explain PCB design decisions.
License
Fine-tuned weights: CC BY-NC-ND 4.0
- Non-commercial use only. You may not use this model or its outputs as part of any commercial product, service, or API.
- No derivatives. You may not fine-tune, merge, quantize, or redistribute modified versions of this model.
- Attribution required. Any publication or project using this model must credit
CanadaDiver/KiCad-Gemma4-E2B-v3.
Base model weights: Apache 2.0 (Google Gemma 4). Both licenses apply.
Commercial licensing: Contact the repository owner for a separate commercial license if you intend to build a product with this model.
Files
| File | Quant | Size | Notes |
|---|---|---|---|
KiCad-Gemma4-E2B-v3-Q8_0.gguf |
Q8_0 | ~5 GB | Full-precision 8-bit — best quality |
Model Details
| Property | Value |
|---|---|
| Base model | google/gemma-4-E2B-it |
| Architecture | Gemma 4 E2B — 35 layers, 4.66 B parameters |
| Attention | Sliding window (256 dim) + global (512 dim), every 5th layer global |
| FFN | Double-wide MLP on global layers (12,288 vs 6,144 for SWA layers) |
| Fine-tuning method | QLoRA via Unsloth |
| LoRA rank / alpha | r=64, α=128, rank-stabilized (rsLoRA) |
| LoRA target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj, embed_tokens, lm_head |
| Context length | 131,072 tokens |
| GGUF quantization | Q8_0 |
| GGUF tensors | 601 (541 base + 60 shared KV projections) |
| Training hardware | NVIDIA GeForce RTX 5080 |
Intended Use
Designed for electronics engineers and hobbyists working with KiCad EDA:
- Generating valid
.kicad_schschematic files from natural language descriptions - Writing KiCad Python scripting API code
- Explaining and debugging KiCad schematics and DRC violations
- Selecting components, writing BOMs
- Creating and modifying footprints and symbols
- Answering KiCad workflow questions
- Agentic schematic generation via the companion MCP server (see below)
Out-of-Scope Use
- General-purpose chat (use the base
gemma-4-E2B-itfor that) - Any commercial product or service without a separate commercial license
- Safety-critical hardware design without qualified human review
Training Data
The model was trained on an extended kicad_agent_final.jsonl dataset — instruction/output pairs assembled from the following sources:
Real Schematics — Seeed Studio OPL KiCad Library
Derived from open-source .kicad_sch and .kicad_sym files published by Seeed Studio.
Boards represented:
- Grove Vision AI Module V2
- Seeed Studio XIAO Series (ESP32S3, nRF52840)
- Wio LR1121 Module v0.9
- Wio SX1262 for XIAO ESP32S3
- Wio SX1262 for XIAO nRF52840 v1.3
- XIAO Family (generic)
License: CC BY-SA 4.0 Source: github.com/Seeed-Studio/OPL_Kicad_Library
These schematics were converted into instruction→output training pairs (schematic generation, section understanding, component reference lookup) using Gemma 4 E4B-it running locally.
Synthetically Generated Pairs
Generated using Gemma 4 E4B-it (via LM Studio) as a teacher model, covering:
| Type | Description |
|---|---|
tool_call_script |
KiCad Python API: generate/modify schematics programmatically |
tool_call_repair |
Debugging and fixing broken KiCad scripts |
sch_generation |
Full schematic generation from circuit description |
sch_section |
Section-level schematic understanding |
Circuit topics covered: passive filters, power supplies (LDO, buck, boost, LiPo charging), microcontroller minimal systems (ESP32, STM32, ATmega, RP2040, nRF52840), digital interfaces (SPI, I2C, UART, USB, CAN, RS-485), analog circuits (op-amps, comparators, ADC/DAC), sensor interfaces, and protection circuits.
Training Procedure
Framework : Unsloth + PEFT + HuggingFace Transformers
Base model : unsloth/gemma-4-E2B-it-unsloth-bnb-4bit (4-bit NF4 for training)
Method : QLoRA (quantized LoRA)
LoRA config:
r = 64
lora_alpha = 128
use_rslora = True (rank-stabilized LoRA)
lora_dropout = 0
use_dora = False
target_modules = q_proj, k_proj, v_proj, o_proj,
gate_proj, up_proj, down_proj,
embed_tokens, lm_head
Training:
Hardware = NVIDIA GeForce RTX 5080 (16 GB)
Sequence length = 4096 tokens max
Train/eval split = 95/5
Framework = SFTTrainer (TRL)
After training, LoRA adapters were merged into the base model weights at bf16 precision and exported to GGUF Q8_0 format via a patched convert_hf_to_gguf.py.
GGUF Compatibility Note
Gemma 4 E2B uses a shared-KV mechanism where layers 15–34 reuse the key/value projections from layers 13 (SWA) and 14 (global). Standard GGUF converters omit these tensors via PyTorch parameter deduplication.
The GGUF file in this repo includes explicit copies of the shared KV tensors for all 35 layers (SWA layers → blk.13 weights, global layers → blk.14 weights), ensuring compatibility with all llama.cpp-based runtimes including LM Studio and Ollama.
Additionally, this file has been patched to correct two issues present in the raw export — see the Technical Notes section below for full details.
Quick Usage
Ollama
# 1. Save a Modelfile
cat > Modelfile <<'EOF'
FROM KiCad-Gemma4-E2B-v3-Q8_0.gguf
TEMPLATE """<bos><start_of_turn>user
{{ .Prompt }}<end_of_turn>
<start_of_turn>model
{{ .Response }}<end_of_turn>
"""
PARAMETER stop "<end_of_turn>"
PARAMETER stop "<start_of_turn>"
PARAMETER temperature 0.2
PARAMETER num_ctx 8192
EOF
# 2. Create the model
ollama create kicad-gemma4-v3 -f Modelfile
⚠️ Important: Use the
/api/chatendpoint with"think": false. Ollama 0.20.x automatically enables Gemma 4 thinking mode which produces empty responses via/api/generate. See Technical Notes below.
import requests
r = requests.post("http://localhost:11434/api/chat", json={
"model": "kicad-gemma4-v3",
"messages": [{"role": "user", "content": "What does DRC stand for in KiCad?"}],
"think": False,
"stream": False
})
print(r.json()["message"]["content"])
LM Studio
- Download
KiCad-Gemma4-E2B-v3-Q8_0.gguf - Place it at:
.lmstudio/models/CanadaDiver/KiCad-Gemma4-E2B-v3/KiCad-Gemma4-E2B-v3-Q8_0.gguf - Select Gemma as the chat template and load — no extra configuration required
Recommended Companion App
For reliable schematic generation in LM Studio, use the companion app:
GalaxyRuler/kicad-gemma-lmstudio-companion
The companion app connects to LM Studio's local API, requests a structured circuit plan from the model, converts that plan into a real KiCad schematic with deterministic Python code, and validates the generated file. Direct raw .kicad_sch generation from plain chat is not the recommended production workflow.
llama.cpp / llama-server
llama-server -m KiCad-Gemma4-E2B-v3-Q8_0.gguf --port 8080 -c 8192
HuggingFace Transformers (full weights)
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "CanadaDiver/KiCad-Gemma4-E2B-v3"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
messages = [{"role": "user", "content": "Write a KiCad schematic for a 3.3V LDO using LP2985."}]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=1024, temperature=1.0, top_k=64, top_p=0.95)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
Agentic Use — KiCad MCP Server
Pair this model with the companion mcp-kicad-sch-api MCP server to give it live read/write access to real .kicad_sch files. The model generates schematic intent; the MCP server executes it against an actual KiCad file on disk.
Install
pip install mcp-kicad-sch-api
Requires Python 3.10+ and KiCad installed (for symbol libraries).
Configure — Claude Desktop
Add to %APPDATA%\Claude\claude_desktop_config.json (Windows) or ~/Library/Application Support/Claude/claude_desktop_config.json (macOS):
{
"mcpServers": {
"kicad-sch-api": {
"command": "python",
"args": ["-m", "mcp_kicad_sch_api"],
"env": {}
}
}
}
Configure — Claude Code
claude mcp add kicad-sch -- python -m mcp_kicad_sch_api
Available MCP Tools
| Tool | Description |
|---|---|
create_schematic |
Create a new .kicad_sch file |
load_schematic |
Load an existing schematic from disk |
save_schematic |
Save the current schematic to disk |
add_component |
Place a component (lib_id, reference, value, position, footprint) |
search_components |
Search KiCad symbol libraries by keyword |
add_wire |
Draw a wire between two coordinates |
add_label |
Add a net label at a position |
add_hierarchical_label |
Add a hierarchical interface label |
add_junction |
Add a junction dot at a wire crossing |
list_components |
List all components in the open schematic |
get_schematic_info |
Get component count, wire count, and file info |
get_component_pin_position |
Get the absolute XY position of a specific pin |
add_label_to_pin |
Attach a net label directly to a component pin |
connect_pins_with_labels |
Connect two pins via a shared net name |
list_component_pins |
List all pins for a component with their positions |
remove_component |
Remove a component by reference |
remove_wire |
Remove a wire by UUID |
Example Agentic Workflow
With the MCP server running alongside this model in Claude Desktop or Claude Code:
"Create a 3.3 V LDO regulator schematic using the LP2985. Add input and output bypass caps. Save it as
ldo_3v3.kicad_sch."
The model will call create_schematic, add_component (×3), add_wire (multiple), add_label, and save_schematic in sequence — producing a valid KiCad 8 schematic file without manual placement.
Hardware & Performance
Tested on RTX 5080 (16 GB VRAM) with ~11 GB already in use:
| Runtime | GPU layers | Speed |
|---|---|---|
| Ollama 0.20.7 | 22–34 / 36 (split offload) | ~19–40 tok/s |
| LM Studio 0.4.11 | full GPU | ~40+ tok/s |
Technical Notes — Bugs Encountered & How They Were Fixed
Getting this GGUF to run correctly under Ollama and LM Studio required resolving three separate bugs, all rooted in how Gemma 4's mixed-architecture interacts with different inference runtimes. Documented here so others do not have to rediscover them.
Bug 1 — Double BOS Token → Empty Output (Ollama)
Symptom: Model loaded cleanly, generated the correct token count at normal speed, but response was always an empty string. eval_count was non-zero, done_reason was stop or length. No error shown.
Root cause: The GGUF metadata field tokenizer.ggml.add_bos_token was set to true. Ollama's tokenizer honoured this and prepended a BOS token (<bos>, ID 2) to every prompt. The Modelfile template also began with <bos>, so the model received a double BOS at position 0. Gemma 4 interprets a double BOS as a malformed sequence and outputs only non-printable special tokens — decoded text is empty.
How it was found: Context token IDs were decoded against the GGUF vocabulary. The first two prompt tokens were both <bos> (ID 2). Ollama logs showed: vocabulary.go:49 warning: adding bos token to prompt which already has it.
Fix: Binary patch of the GGUF — located the tokenizer.ggml.add_bos_token key/value pair (search pattern: uint64-LE string length + key bytes + type code 0x07 for BOOL + value byte 0x01), changed the value byte from 0x01 (true) to 0x00 (false). Ollama no longer prepends BOS automatically; the template's explicit <bos> is the only one present.
Bug 2 — Gemma 4 Thinking Mode Injection → Empty Output (Ollama)
Symptom: Even after fixing Bug 1, template-mode inference still produced an empty response. Raw-mode inference ("raw": true with a fully-formatted prompt) worked correctly and produced fluent answers.
Root cause: Ollama 0.20.x automatically assigns RENDERER gemma4 and PARSER gemma4 to any model whose GGUF architecture is gemma4. The Gemma 4 renderer unconditionally injects a hidden system turn containing the special <|think|> token (ID 98) before the first user turn — regardless of any custom TEMPLATE directive in the Modelfile. This token puts Gemma 4 into chain-of-thought thinking mode. Thinking tokens are filtered from the response field, so with a typical num_predict budget the model never exits the thinking phase and the response is empty.
How it was found: Context token IDs were decoded using the GGUF vocabulary. The decoded prompt sequence was:
<bos> <start_of_turn> system \n <|think|> \n <end_of_turn> \n <start_of_turn> user \n [actual prompt] ...
This hidden system turn was absent in raw mode. The generated tokens decoded to <|channel|> thought \n Thinking Process ... — internal reasoning tokens with no visible text equivalent.
Fix: Use the /api/chat endpoint with "think": false at the request level. The think parameter is not accepted as a Modelfile PARAMETER in Ollama 0.20.x — it must be passed per request. With think: false the system turn is suppressed, prompt token count drops from 31 to 24, and the model responds correctly.
Bug 3 — Tensor Dimension Mismatch → Failed to Load (LM Studio)
Symptom: LM Studio (0.4.11, llama.cpp backend) refused to load the model:
check_tensor_dims: tensor 'blk.15.attn_k.weight' has wrong shape;
expected 1536, 256, got 1536, 512, 1, 1
Root cause: Gemma 4 E2B uses mixed local/global attention. Local (sliding-window) layers have KV head dimension 256; global layers have 512. llama.cpp places global layers at 0-indexed positions 4, 9, 14, 19, 24, 29, 34. The original GGUF was exported with 1-indexed block names (blk.1–blk.35), placing global layers at blk.5, blk.10, blk.15, etc. When llama.cpp read blk.15 it mapped it to 0-indexed layer 15 (a local layer, expecting 256), but found a global-layer tensor (512 + two extra singleton dimensions). This is also why Ollama could load the file — its bundled llama.cpp version did not perform the same strict tensor dimension check.
How it was found: LM Studio server logs at ~/.lmstudio/server-logs/ contained the exact tensor name and expected vs. actual shapes. Cross-referencing with the n_embd_k_gqa array printed at load time confirmed the global/local pattern mismatch.
Fix: The GGUF in this repo was produced by the KV-layer injection pipeline which, as a side effect, re-indexed all block tensors from 1-based to 0-based naming (blk.1→blk.0, …, blk.35→blk.34) and reshaped the global-layer KV tensors to the correct 2D form. The resulting file has blk.0–blk.34 with global layers correctly at 0-indexed positions 4, 9, 14, 19, 24, 29, 34 and all tensor shapes matching llama.cpp's expectations. Both Ollama and LM Studio load this file without errors.
Limitations
- Optimized for KiCad tasks; general reasoning capability is reduced vs. the base model
- Always run KiCad DRC before sending any generated schematic to fabrication
- Trained primarily on KiCad 8.x conventions; KiCad 9.x/10.x syntax may differ in places
- Not a substitute for professional electrical engineering review
Credits & Acknowledgements
| Contributor | Contribution | License |
|---|---|---|
| Google DeepMind | Gemma 4 E2B base model | Apache 2.0 |
| Unsloth | QLoRA fine-tuning framework, 4-bit base | Apache 2.0 |
| Seeed Studio | OPL KiCad Library — real schematic training data | CC BY-SA 4.0 |
| KiCad EDA | KiCad documentation and file format reference | GPL |
| HuggingFace | Transformers, PEFT, TRL training stack | Apache 2.0 |
Synthetic training pairs were generated using Gemma 4 E4B-it running locally via LM Studio.
Citation
@misc{kicad-gemma4-e2b-v3,
author = {CanadaDiver},
title = {KiCad-Gemma4-E2B-v3: A KiCad EDA specialist fine-tuned from Gemma 4 E2B},
year = {2026},
publisher = {HuggingFace},
url = {https://huggingface.co/CanadaDiver/KiCad-Gemma4-E2B-v3-gguf}
}
- Downloads last month
- 196
8-bit