# Mini Whale 1 12B — Local Inference Guide This guide walks you through running Mini Whale 1 12B on a consumer GPU (RTX 3060 12GB or equivalent). ## Prerequisites ### Hardware | Component | Minimum | Recommended | |-----------|---------|-------------| | GPU VRAM | 12 GB | 16 GB+ | | System RAM| 16 GB | 32 GB | | Disk | 25 GB | 50 GB SSD | Tested on: RTX 3060 12GB, 16GB RAM, Windows 11. ### Software ```bash # Python 3.10+ python --version # Create virtual environment python -m venv fuse2-venv # Activate (Windows) fuse2-venv\Scripts\activate # Activate (Linux/Mac) source fuse2-venv/bin/activate # Install dependencies pip install torch --index-url https://download.pytorch.org/whl/cu126 pip install transformers==5.14.1 bitsandbytes==0.49 peft==0.20 accelerate safetensors ``` ## Step 1: Download the Model ```bash pip install huggingface_hub huggingface-cli download Akahsizrr/Mini-Whale-1-12B --local-dir ./mini-whale-1-12b ``` This downloads ~23.4 GB (5 safetensors shards + tokenizer + model code). ## Step 2: Basic Inference (4-bit, ~5 tok/s) This is the simplest way to run the model. All linear layers are quantized to 4-bit NF4. ```python import torch, json, os, sys, gc import torch.nn as nn import torch.nn.functional as F from transformers import AutoConfig, AutoTokenizer from accelerate import init_empty_weights from safetensors import safe_open import bitsandbytes as bnb MODEL_PATH = "./mini-whale-1-12b" DEVICE = "cuda:0" sys.path.insert(0, MODEL_PATH) import fuse2_model_local # ── Step 2a: Load config ── config = AutoConfig.from_pretrained(MODEL_PATH, trust_remote_code=True) config._attn_implementation = "sdpa" # Critical: use SDPA, not eager # ── Step 2b: Create model on meta device (saves RAM) ── with init_empty_weights(): model = fuse2_model_local.Fuse2ForCausalLM(config) # ── Step 2c: Load weights with 4-bit quantization ── quant_suffixes = ("q_proj.weight", "k_proj.weight", "v_proj.weight", "o_proj.weight", "gate_proj.weight", "up_proj.weight", "down_proj.weight") with open(f"{MODEL_PATH}/model.safetensors.index.json") as f: index = json.load(f) weight_map = index["weight_map"] def nav(model, key): parts = key.split(".") obj = model for p in parts[:-1]: obj = obj[int(p)] if p.isdigit() else getattr(obj, p) return obj, parts[-1] def find_linear(model, key): parts = key.split(".") obj = model for p in parts[:-2]: obj = obj[int(p)] if p.isdigit() else getattr(obj, p) return obj, parts[-2] param_names = set(dict(model.named_parameters()).keys()) replaced = {} for shard_name in sorted(set(weight_map.values())): with safe_open(os.path.join(MODEL_PATH, shard_name), framework="pt", device="cpu") as f: for key in [k for k, v in weight_map.items() if v == shard_name]: if key not in param_names: continue tensor = f.get_tensor(key) if any(key.endswith(s) for s in quant_suffixes): owner, attr = find_linear(model, key) old = getattr(owner, attr) new = bnb.nn.Linear4bit(old.in_features, old.out_features, bias=False, quant_type="nf4", compute_dtype=torch.bfloat16, device=DEVICE) new.weight = bnb.nn.Params4bit(tensor.to(torch.bfloat16), requires_grad=False, quant_type="nf4").cuda(0) setattr(owner, attr, new) else: parent, pname = nav(model, key) parent._parameters[pname] = nn.Parameter( tensor.to(torch.bfloat16).to(DEVICE), requires_grad=False) del tensor gc.collect() torch.cuda.empty_cache() # ── Step 2d: Fix tied embeddings ── if model.lm_head.weight.device.type == 'meta': model.lm_head.weight = nn.Parameter( model.model.embed_tokens.weight.data.clone(), requires_grad=False) # ── Step 2e: Apply runtime fixes (SwiGLU clamp + router stability) ── from fuse2_model_local import Fuse2AugmentedLayer for layer in model.model.layers: if not isinstance(layer, Fuse2AugmentedLayer): continue if hasattr(layer, 'coding_gate') and layer.coding_gate.device.type == 'meta': layer.coding_gate = nn.Parameter(torch.tensor(-2.0, device=DEVICE)) if hasattr(layer, 'coding_norm') and layer.coding_norm.weight.device.type == 'meta': layer.coding_norm = nn.RMSNorm( layer.coding_norm.weight.shape[0], eps=1e-6).to(DEVICE) experts = getattr(layer, "experts", None) if experts: for expert in experts: gp, up, dp = expert.gate_proj, expert.up_proj, expert.down_proj def make_fwd(g, u, d, lim=10.0): def forward(x): return d(torch.clamp(F.silu(g(x)) * u(x), -lim, lim)) return forward expert.forward = make_fwd(gp, up, dp) router = getattr(layer, "router", None) if router: gate, top_k = router.gate, router.top_k def make_router(g, tk): def forward(h): logits = g(h) scores = torch.clamp(F.softplus(logits), min=1e-6).sqrt() w, idx = scores.topk(tk, dim=-1) return w / (w.sum(dim=-1, keepdim=True) + 1e-8), idx, logits return forward router.forward = make_router(gate, top_k) # ── Step 2f: Finalize model ── model.set_coding_enabled(True) model.to(DEVICE) model.eval() # Fix norm dtypes (float32 → bfloat16 for fused kernels) for module in model.modules(): if hasattr(module, 'weight') and hasattr(module, 'eps'): if module.weight.dtype == torch.float32: module.weight.data = module.weight.data.to(torch.bfloat16) print(f"VRAM: {torch.cuda.memory_allocated(0)/1e9:.1f} GB") # ── Step 2g: Generate ── tok = AutoTokenizer.from_pretrained(MODEL_PATH, trust_remote_code=True) messages = [{"role": "user", "content": "Write a Python fizzbuzz."}] text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) input_ids = tok(text, return_tensors="pt").input_ids.to(DEVICE) with torch.inference_mode(): out = model.generate(input_ids, max_new_tokens=512, do_sample=False, repetition_penalty=1.3) print(tok.decode(out[0], skip_special_tokens=True)) ``` **Expected output:** ``` VRAM: 8.9 GB Okay, I need to write a Python fizzbuzz. Let me think about this... ``` ## Step 3: DSpark Speculative Decoding (~10 tok/s) For 2x speedup, use the included `fuse2_dspark_fast.py` script: ```bash python fuse2_dspark_fast.py "Write a Python function to check if a number is prime." 512 ``` This requires the DSpark drafter model. The drafter is a 5-layer Qwen3 model that predicts 7 tokens per block. ### How DSpark works ``` ┌─────────────┐ ┌──────────────┐ │ Drafter │────▶│ Target │ │ (5 layers) │ │ (36 layers) │ │ Predicts │ │ Verifies │ │ 7 tokens │ │ in 1 pass │ └─────────────┘ └──────────────┘ ~32ms ~330ms ``` 1. Drafter predicts 7 tokens using target's hidden states 2. Target verifies all 7 in a single forward pass 3. Accept the longest matching prefix (greedy argmax) 4. Target produces 1 bonus token 5. Repeat **Acceptance rate**: 2.3-3.4/7 (33-49%), giving 1.8x speedup on average. ### DSpark VRAM optimization The drafter is quantized to 4-bit and shares the embedding layer with the target: | Component | VRAM | |-----------|------| | Target (BF16 attention + 4-bit MLP) | 10.3 GB | | Drafter (4-bit, shared embedding) | 0.6 GB | | KV cache + activations | ~0.5 GB | | **Total** | **~10.9 GB (fits in 12 GB)** | ## Step 4: Sliding Window KV Cache For long sequences (512+ tokens), the KV cache grows and slows down generation. The DSpark script includes a sliding window that limits the KV cache to 512 tokens: ```python MAX_KV_CACHE = 512 # tokens cache_len = target_cache.get_seq_length() if cache_len > MAX_KV_CACHE: crop_to = cache_len - MAX_KV_CACHE for layer in target_cache.layers: if hasattr(layer, 'keys') and layer.keys.numel() > 0: layer.keys = layer.keys[:, :, crop_to:, :] layer.values = layer.values[:, :, crop_to:, :] ``` This keeps speed constant at ~10 tok/s regardless of sequence length. ## Troubleshooting ### "CUDA out of memory" - Check `torch.cuda.memory_reserved(0)` — if it's >11.5 GB on a 12 GB card, you're spilling to system RAM - Make sure ALL linear layers are 4-bit (including attention) - Use `gc.collect(); torch.cuda.empty_cache()` after loading each shard - Reduce `max_new_tokens` to 256 ### "TritonMissing" error - `torch.compile` is not supported on Windows (no triton) - The model code already disables this — make sure you're using the latest `fuse2_model_local.py` ### "expandable_segments not supported" - This is a harmless warning on Windows - The model works fine without it ### Generation is garbage / repetitive - Use `do_sample=False` (greedy decoding) — sampling can produce NaN - Use `repetition_penalty=1.3` - Make sure `coding_enabled=True` (the experts need to be on) - Check that SwiGLU clamping is applied (limit=10.0) ### Speed is 1-2 tok/s instead of 10 - **This means VRAM is overflowing to system RAM** - Check `torch.cuda.memory_reserved(0)` — it should be <11.5 GB - Switch to 4-bit quantization for ALL layers (including attention) - Quantize the drafter to 4-bit - Share the embedding between target and drafter ## Performance Summary | Mode | Speed | VRAM | Best for | |------|-------|------|----------| | Basic 4-bit | 5.5 tok/s | 8.9 GB | Simple tasks, max VRAM headroom | | DSpark (4-bit all) | 8.5 tok/s | 10.3 GB | Balanced | | DSpark (BF16 attn + shared) | 10.0 tok/s | 10.9 GB | **Best speed** |