whoy commited on
Commit
b0d1386
·
verified ·
1 Parent(s): a040386

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +126 -3
README.md CHANGED
@@ -1,3 +1,126 @@
1
- ---
2
- license: mit
3
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ru
4
+ - en
5
+ license: mit
6
+ base_model: ai-sage/GigaChat3-10B-A1.8B-bf16
7
+ tags:
8
+ - gguf
9
+ - llama.cpp
10
+ - experimental
11
+ - unstable
12
+ - moe
13
+ model_type: deepseek_v3
14
+ library_name: llama.cpp
15
+ ---
16
+
17
+ # GigaChat3-10B-A1.8B GGUF [EXPERIMENTAL]
18
+
19
+ ⚠️ **UNSTABLE BUILD** — This is an experimental GGUF conversion with known quality issues. Use for testing only.
20
+
21
+ ## What is this?
22
+
23
+ Experimental GGUF conversion of [GigaChat3-10B-A1.8B](https://huggingface.co/ai-sage/GigaChat3-10B-A1.8B) — a Russian dialogue model with MoE + MLA architecture.
24
+
25
+ **Model specs:**
26
+ - 10B parameters (1.8B active)
27
+ - 64 experts, 4 active per token
28
+ - 262k context window
29
+ - BF16 → GGUF conversion
30
+
31
+ ## ⚠️ Known Issues
32
+
33
+ **This conversion has degraded quality compared to the original model** due to architectural incompatibility:
34
+
35
+ 1. **Hybrid MLA problem:** GigaChat3 uses standard Q-projection (no compression) + compressed KV-cache, which llama.cpp doesn't support natively
36
+ 2. **RoPE mismatch:** Position embeddings are applied in wrong dimensional space
37
+ 3. **Symptoms:** Incoherent long-form generation, context confusion, occasional nonsense
38
+
39
+ **Why it still loads:** We emulated missing MLA components using Identity matrices, which satisfies llama.cpp's loader but breaks positional logic.
40
+
41
+ ## When to use this
42
+
43
+ ✅ **Good for:**
44
+ - Short prompts (1-3 turns)
45
+ - Fact retrieval / memorized knowledge
46
+ - Testing GGUF tooling compatibility
47
+ - Placeholder until proper support arrives
48
+
49
+ ❌ **Bad for:**
50
+ - Production use
51
+ - Long conversations
52
+ - Complex reasoning tasks
53
+ - Anything requiring positional awareness
54
+
55
+ ## Conversion method
56
+
57
+ ```python
58
+ # 1. Restructure weights to emulate MLA
59
+ # Original: Q = X @ q_proj [6144, 1536]
60
+ # Emulated: Q = ((X @ Identity[1536,1536]) * ones) @ q_proj[6144,1536]
61
+
62
+ # 2. Convert with q_lora_rank = 1536
63
+ python prepare_weights.py # Creates fake q_a_proj, q_a_norm, q_b_proj
64
+ python convert_hf_to_gguf.py ./model-fixed --outfile model.gguf
65
+ ```
66
+
67
+ **Math is preserved, but RoPE positioning is broken.**
68
+
69
+ ## Usage
70
+
71
+ ```bash
72
+ # llama.cpp
73
+ ./llama-cli -m model.gguf \
74
+ --temp 0.3 --top-p 0.9 -n 512 \
75
+ -p "User: [query]\nAssistant:"
76
+
77
+ # Recommended params
78
+ temperature: 0.0-0.5
79
+ top_p: 0.8-0.9
80
+ max_tokens: < 512 (quality degrades further out)
81
+ ```
82
+
83
+ ## Chat template
84
+
85
+ Use this in LM Studio or add to tokenizer_config.json:
86
+
87
+ ```
88
+ developer system<|role_sep|>
89
+ You are a helpful assistant.
90
+ <|message_sep|>
91
+
92
+ user<|role_sep|>
93
+ {prompt}
94
+ <|message_sep|>
95
+
96
+ assistant<|role_sep|>
97
+ ```
98
+
99
+ ## Better alternatives
100
+
101
+ For production quality, use the original model with:
102
+ - **vLLM** (native FP8 support, proper inference)
103
+ - **transformers** (HF native, slower but correct)
104
+ - **SGLang** (fast + correct)
105
+
106
+ Or wait for proper llama.cpp support (requires C++ patch).
107
+
108
+ ## Technical details
109
+
110
+ **Problem:** llama.cpp DeepSeek implementation assumes Q-vectors are compressed (q_lora_rank < hidden_size). GigaChat3 skips Q-compression.
111
+
112
+ **Hack:** Set q_lora_rank = hidden_size (1536) and inject Identity matrices to fake compression.
113
+
114
+ **Result:** Loader accepts it, but RoPE gets applied to wrong intermediate representation → broken positional encoding → quality loss.
115
+
116
+ ## Future
117
+
118
+ If you're a llama.cpp dev: The fix is adding a branch for `q_lora_rank == null` in the DeepSeek V3 \ V2 attention code (~100 LOC). Happy to help test!
119
+
120
+ ## License
121
+
122
+ MIT (inherited from base model)
123
+
124
+ ---
125
+
126
+ **TL;DR:** Works technically, quality is compromised. Use vLLM for real work.