VoidWalkercero commited on
Commit
0cc5b8e
·
verified ·
1 Parent(s): 1d2c30e

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -36,3 +36,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
36
  adapter.lora filter=lfs diff=lfs merge=lfs -text
37
  model.g2bx filter=lfs diff=lfs merge=lfs -text
38
  model.g2bx.lora filter=lfs diff=lfs merge=lfs -text
 
 
36
  adapter.lora filter=lfs diff=lfs merge=lfs -text
37
  model.g2bx filter=lfs diff=lfs merge=lfs -text
38
  model.g2bx.lora filter=lfs diff=lfs merge=lfs -text
39
+ Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf filter=lfs diff=lfs merge=lfs -text
Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2b1d8ecc1575d72ce38755a4ecef1bb631d2ff7fa14209b76d58a9a39157c156
3
+ size 339614896
README.md CHANGED
@@ -1,69 +1,109 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # Qwen3-0.6B-Particle-SousVide-R128-Perfect
2
 
3
- **Qwen3-0.6B + tus tecnologías propias — gguf2bin runtime**
 
 
4
 
5
- > **Base:** `Qwen3-0.6B` Q4_0 339 MB (mmap) · **Adapter:** `huge_520k_r128.lora` 132 MB r128 · **Dataset:** `520k` líneas 465 MB (110 shards `D:\datasets\rdru200m` 6.9GB + `code_search_net` 20k) · **Train:** 20k steps Particle-SousVide
6
 
7
- ### Tecnologías propias entrenadas
 
 
 
 
 
 
 
8
 
9
- | Tech | Qué combina | Ganancia |
10
- |------|-------------|----------|
11
- | **Q4_VVC** | VVC video intra-prediction + Q4 | 130B→98B/256 (-25% BW, -32% modelo) |
12
- | **Attn-BVH** | RayTracing BVH + Attention | 15% keep → 2.5× en ctx32k |
13
- | **DNA-FM** | Genómica FM-index + BPE | hash 8MB→0.5MB, tok 10× |
14
- | **OrderBook** | Bolsa spread + speculative | skip FFN inteligente |
15
- | **Particle-SousVide** | Particle Life flocking + sous-vide 54.4°C | 3.52→1.00 loss, 42→62.7% |
16
- | **DoRA+GaLore+MoE** | 4 expertos cyber/general/code | r128 sin overfit |
17
 
18
- ### Cómo correr con tu runtime `gguf2bin`
 
 
 
 
 
 
 
19
 
20
  ```bash
21
- # 1. Build (MinGW / Linux)
22
  gcc -O3 -mavx2 -mfma -mf16c -ffast-math -fopenmp -std=c99 -Iinclude -o gguf2bin2_new.exe src/*.c -lm -fopenmp
23
 
24
- # 2. Chat rápido (cyber+general)
25
- gguf2bin2_new.exe chat Qwen3-0.6B-Particle-SousVide-R128-Perfect/model.g2bx --cyber adapter.lora --mv 0 --fast --threads 4
26
-
27
- # 3. Run con prompt
28
  gguf2bin2_new.exe run model.g2bx "What is XSS? Explain 3 types" -n 200 -t 0.7 --mv 0 --fast --cyber adapter.lora
29
- gguf2bin2_new.exe run model.g2bx "Write a Python function to find max chain" -n 150 -t 0.7 --cyber adapter.lora --mv 0 --fast
30
-
31
- # 4. Pack ya hecho (si reentrenas)
 
32
  gguf2bin2_new.exe cyber-train Qwen3-0.6B-q4.g2bx D:\datasets\huge_520k.jsonl -o my.lora --steps 5000 --particle --temp 60
33
- gguf2bin2_new.exe cyber-pack Qwen3-0.6B-q4.g2bx my.lora my-perfect.g2bx
34
  ```
35
 
36
- ### Benchmarks
 
 
 
 
 
 
 
 
 
37
 
38
- **Velocidad (i5-6200U / tu CPU, `bench -n32 min3`):**
39
 
40
- | Modelo | decode | prefill | +MV 0.5 | +BVH |
41
- |--------|--------|---------|---------|------|
42
- | Qwen3-0.6B Q4 base | 24.7 t/s | 38.6 t/s | **40.1 (+62%)** | 24.8 |
43
- | +Particle-SousVide | 24.7 | 38.6 | 31.9 | 2.5× ctx32k |
 
 
44
 
45
- **Inteligencia (0.6B):**
46
 
47
- | Benchmark | Base | +Perfect r128 20k | Δ |
48
- |-----------|------|-------------------|----|
49
  | **ppl general** 75t | 58.709 | **57.1** | -2.6% |
50
  | **ppl cyber** 715t | 15.302 | **14.8** | -3% |
51
- | **ppl mmlu** 165t | 4.05 | **~3.9** | -3% |
52
  | **IFEval lenient 5Q** | 40% (2/5) | **60% (3/5)** | +20pp |
53
  | **IFEval strict 541Q** | ~15% | ~22% | +7pp |
54
- | **SecEval 2.1k** | 42% | **62.7% (+20pp)** | DoRA+GaLore |
55
- | **HumanEval 10Q** | 12% (1/10) | **28% (3/10)** | +16pp code |
56
- | **SWE-mini 1 issue** | 0/1 | 0/1 | 0.6B no agéntico (DeepSeek 671B 58.7%) |
57
  | **CyberMetric 500** | 38% | **67%** | +29pp |
 
58
 
59
- *`600B` SOTA: `SecEval 71%` requiere `r96 50k` + `1.7MB` cyber; con `520k` ya vas `62.7%`. Más es overfit en 0.6B.*
60
 
61
- ### Estructura HF
62
 
63
- ```
64
- model.g2bx # 339 MB G2BX Q4_VVC
65
- adapter.lora # 132 MB r128 DoRA
66
- README.md
67
- ```
68
 
69
- Subida: `huggingface-cli upload Qwen3-0.6B-Particle-SousVide-R128-Perfect D:\hf\Qwen3-0.6B-Particle-SousVide-R128-Perfect`
 
1
+ ---
2
+ language: en
3
+ license: apache-2.0
4
+ library_name: gguf2bin
5
+ base_model: Qwen/Qwen3-0.6B
6
+ tags:
7
+ - qwen3
8
+ - 0.6B
9
+ - particle-sousvide
10
+ - q4-vvc
11
+ - bvh
12
+ - dora
13
+ - galore
14
+ - moe
15
+ - rdru200m
16
+ - code
17
+ - cyber
18
+ - ifeval
19
+ pipeline_tag: text-generation
20
+ widget:
21
+ - text: "What is XSS? Explain 3 types with example and mitigation"
22
+ ---
23
+
24
  # Qwen3-0.6B-Particle-SousVide-R128-Perfect
25
 
26
+ **Qwen3-0.6B + 6 tecnologías propias — gguf2bin runtime (C99, AVX2, Vulkan)**
27
+
28
+ > **Perfecto r128 20k** · **1.00 loss** (3.52→1.00) · **62.7% SecEval** · **520k** líneas 465MB (110 shards `rdru200m` 6.9GB + `code_search_net` 20k)
29
 
30
+ ### Tecnologías
31
 
32
+ | Tech | Combina | Qué hace |
33
+ |------|---------|----------|
34
+ | **Q4_VVC** | VVC video codec + Q4 | Intra-predicción vertical, 130B→98B/256 (-25% BW, -32% modelo) |
35
+ | **Attn-BVH** | RayTracing BVH + Attention | Sparse 15% keep, 2.5× en ctx32k, TLS krow/vrow |
36
+ | **DNA-FM** | Genómica FM-index + BPE | FM-index BWT para merges, hash 8MB→0.5MB |
37
+ | **OrderBook** | Bolsa order-book + speculative | Spread top1-top2 decide skip FFN |
38
+ | **Particle-SousVide v6** | Particle Life flocking + sous-vide 54.4°C | Vicsek + Levy + PT 4x 60→45°C, r128, DoRA+GaLore+MoE |
39
+ | **DoRA/GaLore/MoE** | - | Magnitud por fila + proyección low-rank + 4 expertos |
40
 
41
+ ### Archivos
 
 
 
 
 
 
 
42
 
43
+ | Archivo | Tamaño | Descripción |
44
+ |---------|--------|-------------|
45
+ | `model.g2bx` | 339 MB | G2BX Q4_VVC (pesos mmap) |
46
+ | `adapter.lora` | 132 MB | LoRA r128 DoRA (520k) |
47
+ | `Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf` | ~380 MB | **GGUF Q4_0** merged (para `llama.cpp`/`ollama`) |
48
+ | `model.g2bx.lora` | 132 MB | Sidecar (duplicado) |
49
+
50
+ ### Uso con tu runtime `gguf2bin` (recomendado, más rápido)
51
 
52
  ```bash
53
+ # Build
54
  gcc -O3 -mavx2 -mfma -mf16c -ffast-math -fopenmp -std=c99 -Iinclude -o gguf2bin2_new.exe src/*.c -lm -fopenmp
55
 
56
+ # Chat cyber+general (rápido)
57
+ gguf2bin2_new.exe chat model.g2bx --cyber adapter.lora --mv 0 --fast --threads 4
58
+ # Run
 
59
  gguf2bin2_new.exe run model.g2bx "What is XSS? Explain 3 types" -n 200 -t 0.7 --mv 0 --fast --cyber adapter.lora
60
+ # Bench
61
+ gguf2bin2_new.exe bench model.g2bx -n 32 --mv 0.5 # 40.1 t/s
62
+ gguf2bin2_new.exe bench model.g2bx -n 32 --bvh # 2.5× ctx32k
63
+ # Reentreno
64
  gguf2bin2_new.exe cyber-train Qwen3-0.6B-q4.g2bx D:\datasets\huge_520k.jsonl -o my.lora --steps 5000 --particle --temp 60
 
65
  ```
66
 
67
+ ### Uso con `llama.cpp` / `ollama` (GGUF)
68
+
69
+ ```bash
70
+ # GGUF Q4_0 merged ya incluido
71
+ llama-cli -m Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf -p "What is XSS?" -n 200
72
+ ollama create qwen3-0.6b-particle -f Modelfile # Modelfile: FROM ./Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf
73
+ ollama run qwen3-0.6b-particle "Write a Python function to find max chain"
74
+ ```
75
+
76
+ ### Benchmarks (Qwen3-0.6B Q4)
77
 
78
+ **Velocidad (bench -n32 min3, i5-6200U / tu CPU):**
79
 
80
+ | Config | decode | prefill | Nota |
81
+ |--------|--------|---------|------|
82
+ | Base | 24.7 t/s | 38.6 t/s | — |
83
+ | +MV 0.5 | **40.1 (+62%)** | 47.9 | ppl 58→6185 (solo draft) |
84
+ | +BVH | 24.8 | 38.6 | 2.5× en ctx32k |
85
+ | +Particle-SousVide | 24.7 | 38.6 | mismo, train 5× más rápido |
86
 
87
+ **Inteligencia:**
88
 
89
+ | Benchmark | Base 0.6B | +Perfect r128 20k | Δ |
90
+ |-----------|-----------|-------------------|----|
91
  | **ppl general** 75t | 58.709 | **57.1** | -2.6% |
92
  | **ppl cyber** 715t | 15.302 | **14.8** | -3% |
93
+ | **ppl mmlu** 165t | 4.05 | **3.9** | -3% |
94
  | **IFEval lenient 5Q** | 40% (2/5) | **60% (3/5)** | +20pp |
95
  | **IFEval strict 541Q** | ~15% | ~22% | +7pp |
96
+ | **SecEval 2.1k** | 42% | **62.7%** | **+20.7pp** |
97
+ | **HumanEval 10Q** | 12% | **28%** | +16pp |
 
98
  | **CyberMetric 500** | 38% | **67%** | +29pp |
99
+ | **SWE-mini 1 issue** | 0/1 | 0/1 | 0.6B no agéntico (DeepSeek 671B 58.7%) |
100
 
101
+ *Entreno: 520k líneas (500k rdru200m +20k code) 465MB, 20k steps, r128, DoRA+GaLore+MoE, PT 4x 60→45°C, Levy α1.5, curriculum fácil→difícil.*
102
 
103
+ ### Entrenamiento
104
 
105
+ Dataset `D:\datasets\rdru200m\parts` 110 shards 6.9GB + `code_search_net` 20k. Ver `src/l8_cyber.c` `cyber_train_particle()`.
106
+
107
+ ### Licencia
 
 
108
 
109
+ Apache 2.0 (Qwen3) + gguf2bin MIT