Upload folder using huggingface_hub
Browse files- .gitattributes +1 -0
- Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf +3 -0
- README.md +82 -42
.gitattributes
CHANGED
|
@@ -36,3 +36,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 36 |
adapter.lora filter=lfs diff=lfs merge=lfs -text
|
| 37 |
model.g2bx filter=lfs diff=lfs merge=lfs -text
|
| 38 |
model.g2bx.lora filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 36 |
adapter.lora filter=lfs diff=lfs merge=lfs -text
|
| 37 |
model.g2bx filter=lfs diff=lfs merge=lfs -text
|
| 38 |
model.g2bx.lora filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf filter=lfs diff=lfs merge=lfs -text
|
Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:2b1d8ecc1575d72ce38755a4ecef1bb631d2ff7fa14209b76d58a9a39157c156
|
| 3 |
+
size 339614896
|
README.md
CHANGED
|
@@ -1,69 +1,109 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Qwen3-0.6B-Particle-SousVide-R128-Perfect
|
| 2 |
|
| 3 |
-
**Qwen3-0.6B +
|
|
|
|
|
|
|
| 4 |
|
| 5 |
-
|
| 6 |
|
| 7 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|
| 9 |
-
|
| 10 |
-
|------|-------------|----------|
|
| 11 |
-
| **Q4_VVC** | VVC video intra-prediction + Q4 | 130B→98B/256 (-25% BW, -32% modelo) |
|
| 12 |
-
| **Attn-BVH** | RayTracing BVH + Attention | 15% keep → 2.5× en ctx32k |
|
| 13 |
-
| **DNA-FM** | Genómica FM-index + BPE | hash 8MB→0.5MB, tok 10× |
|
| 14 |
-
| **OrderBook** | Bolsa spread + speculative | skip FFN inteligente |
|
| 15 |
-
| **Particle-SousVide** | Particle Life flocking + sous-vide 54.4°C | 3.52→1.00 loss, 42→62.7% |
|
| 16 |
-
| **DoRA+GaLore+MoE** | 4 expertos cyber/general/code | r128 sin overfit |
|
| 17 |
|
| 18 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
|
| 20 |
```bash
|
| 21 |
-
#
|
| 22 |
gcc -O3 -mavx2 -mfma -mf16c -ffast-math -fopenmp -std=c99 -Iinclude -o gguf2bin2_new.exe src/*.c -lm -fopenmp
|
| 23 |
|
| 24 |
-
#
|
| 25 |
-
gguf2bin2_new.exe chat
|
| 26 |
-
|
| 27 |
-
# 3. Run con prompt
|
| 28 |
gguf2bin2_new.exe run model.g2bx "What is XSS? Explain 3 types" -n 200 -t 0.7 --mv 0 --fast --cyber adapter.lora
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
|
|
|
| 32 |
gguf2bin2_new.exe cyber-train Qwen3-0.6B-q4.g2bx D:\datasets\huge_520k.jsonl -o my.lora --steps 5000 --particle --temp 60
|
| 33 |
-
gguf2bin2_new.exe cyber-pack Qwen3-0.6B-q4.g2bx my.lora my-perfect.g2bx
|
| 34 |
```
|
| 35 |
|
| 36 |
-
###
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 37 |
|
| 38 |
-
**Velocidad (i5-6200U / tu CPU
|
| 39 |
|
| 40 |
-
|
|
| 41 |
-
|--------|--------|---------|------
|
| 42 |
-
|
|
| 43 |
-
| +
|
|
|
|
|
|
|
| 44 |
|
| 45 |
-
**Inteligencia
|
| 46 |
|
| 47 |
-
| Benchmark | Base | +Perfect r128 20k | Δ |
|
| 48 |
-
|-----------|------|-------------------|----|
|
| 49 |
| **ppl general** 75t | 58.709 | **57.1** | -2.6% |
|
| 50 |
| **ppl cyber** 715t | 15.302 | **14.8** | -3% |
|
| 51 |
-
| **ppl mmlu** 165t | 4.05 | **
|
| 52 |
| **IFEval lenient 5Q** | 40% (2/5) | **60% (3/5)** | +20pp |
|
| 53 |
| **IFEval strict 541Q** | ~15% | ~22% | +7pp |
|
| 54 |
-
| **SecEval 2.1k** | 42% | **62.7%
|
| 55 |
-
| **HumanEval 10Q** | 12%
|
| 56 |
-
| **SWE-mini 1 issue** | 0/1 | 0/1 | 0.6B no agéntico (DeepSeek 671B 58.7%) |
|
| 57 |
| **CyberMetric 500** | 38% | **67%** | +29pp |
|
|
|
|
| 58 |
|
| 59 |
-
*
|
| 60 |
|
| 61 |
-
###
|
| 62 |
|
| 63 |
-
```
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
README.md
|
| 67 |
-
```
|
| 68 |
|
| 69 |
-
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: en
|
| 3 |
+
license: apache-2.0
|
| 4 |
+
library_name: gguf2bin
|
| 5 |
+
base_model: Qwen/Qwen3-0.6B
|
| 6 |
+
tags:
|
| 7 |
+
- qwen3
|
| 8 |
+
- 0.6B
|
| 9 |
+
- particle-sousvide
|
| 10 |
+
- q4-vvc
|
| 11 |
+
- bvh
|
| 12 |
+
- dora
|
| 13 |
+
- galore
|
| 14 |
+
- moe
|
| 15 |
+
- rdru200m
|
| 16 |
+
- code
|
| 17 |
+
- cyber
|
| 18 |
+
- ifeval
|
| 19 |
+
pipeline_tag: text-generation
|
| 20 |
+
widget:
|
| 21 |
+
- text: "What is XSS? Explain 3 types with example and mitigation"
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
# Qwen3-0.6B-Particle-SousVide-R128-Perfect
|
| 25 |
|
| 26 |
+
**Qwen3-0.6B + 6 tecnologías propias — gguf2bin runtime (C99, AVX2, Vulkan)**
|
| 27 |
+
|
| 28 |
+
> **Perfecto r128 20k** · **1.00 loss** (3.52→1.00) · **62.7% SecEval** · **520k** líneas 465MB (110 shards `rdru200m` 6.9GB + `code_search_net` 20k)
|
| 29 |
|
| 30 |
+
### Tecnologías
|
| 31 |
|
| 32 |
+
| Tech | Combina | Qué hace |
|
| 33 |
+
|------|---------|----------|
|
| 34 |
+
| **Q4_VVC** | VVC video codec + Q4 | Intra-predicción vertical, 130B→98B/256 (-25% BW, -32% modelo) |
|
| 35 |
+
| **Attn-BVH** | RayTracing BVH + Attention | Sparse 15% keep, 2.5× en ctx32k, TLS krow/vrow |
|
| 36 |
+
| **DNA-FM** | Genómica FM-index + BPE | FM-index BWT para merges, hash 8MB→0.5MB |
|
| 37 |
+
| **OrderBook** | Bolsa order-book + speculative | Spread top1-top2 decide skip FFN |
|
| 38 |
+
| **Particle-SousVide v6** | Particle Life flocking + sous-vide 54.4°C | Vicsek + Levy + PT 4x 60→45°C, r128, DoRA+GaLore+MoE |
|
| 39 |
+
| **DoRA/GaLore/MoE** | - | Magnitud por fila + proyección low-rank + 4 expertos |
|
| 40 |
|
| 41 |
+
### Archivos
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
|
| 43 |
+
| Archivo | Tamaño | Descripción |
|
| 44 |
+
|---------|--------|-------------|
|
| 45 |
+
| `model.g2bx` | 339 MB | G2BX Q4_VVC (pesos mmap) |
|
| 46 |
+
| `adapter.lora` | 132 MB | LoRA r128 DoRA (520k) |
|
| 47 |
+
| `Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf` | ~380 MB | **GGUF Q4_0** merged (para `llama.cpp`/`ollama`) |
|
| 48 |
+
| `model.g2bx.lora` | 132 MB | Sidecar (duplicado) |
|
| 49 |
+
|
| 50 |
+
### Uso con tu runtime `gguf2bin` (recomendado, más rápido)
|
| 51 |
|
| 52 |
```bash
|
| 53 |
+
# Build
|
| 54 |
gcc -O3 -mavx2 -mfma -mf16c -ffast-math -fopenmp -std=c99 -Iinclude -o gguf2bin2_new.exe src/*.c -lm -fopenmp
|
| 55 |
|
| 56 |
+
# Chat cyber+general (rápido)
|
| 57 |
+
gguf2bin2_new.exe chat model.g2bx --cyber adapter.lora --mv 0 --fast --threads 4
|
| 58 |
+
# Run
|
|
|
|
| 59 |
gguf2bin2_new.exe run model.g2bx "What is XSS? Explain 3 types" -n 200 -t 0.7 --mv 0 --fast --cyber adapter.lora
|
| 60 |
+
# Bench
|
| 61 |
+
gguf2bin2_new.exe bench model.g2bx -n 32 --mv 0.5 # 40.1 t/s
|
| 62 |
+
gguf2bin2_new.exe bench model.g2bx -n 32 --bvh # 2.5× ctx32k
|
| 63 |
+
# Reentreno
|
| 64 |
gguf2bin2_new.exe cyber-train Qwen3-0.6B-q4.g2bx D:\datasets\huge_520k.jsonl -o my.lora --steps 5000 --particle --temp 60
|
|
|
|
| 65 |
```
|
| 66 |
|
| 67 |
+
### Uso con `llama.cpp` / `ollama` (GGUF)
|
| 68 |
+
|
| 69 |
+
```bash
|
| 70 |
+
# GGUF Q4_0 merged ya incluido
|
| 71 |
+
llama-cli -m Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf -p "What is XSS?" -n 200
|
| 72 |
+
ollama create qwen3-0.6b-particle -f Modelfile # Modelfile: FROM ./Qwen3-0.6B-Particle-SousVide-R128-Perfect.gguf
|
| 73 |
+
ollama run qwen3-0.6b-particle "Write a Python function to find max chain"
|
| 74 |
+
```
|
| 75 |
+
|
| 76 |
+
### Benchmarks (Qwen3-0.6B Q4)
|
| 77 |
|
| 78 |
+
**Velocidad (bench -n32 min3, i5-6200U / tu CPU):**
|
| 79 |
|
| 80 |
+
| Config | decode | prefill | Nota |
|
| 81 |
+
|--------|--------|---------|------|
|
| 82 |
+
| Base | 24.7 t/s | 38.6 t/s | — |
|
| 83 |
+
| +MV 0.5 | **40.1 (+62%)** | 47.9 | ppl 58→6185 (solo draft) |
|
| 84 |
+
| +BVH | 24.8 | 38.6 | 2.5× en ctx32k |
|
| 85 |
+
| +Particle-SousVide | 24.7 | 38.6 | mismo, train 5× más rápido |
|
| 86 |
|
| 87 |
+
**Inteligencia:**
|
| 88 |
|
| 89 |
+
| Benchmark | Base 0.6B | +Perfect r128 20k | Δ |
|
| 90 |
+
|-----------|-----------|-------------------|----|
|
| 91 |
| **ppl general** 75t | 58.709 | **57.1** | -2.6% |
|
| 92 |
| **ppl cyber** 715t | 15.302 | **14.8** | -3% |
|
| 93 |
+
| **ppl mmlu** 165t | 4.05 | **3.9** | -3% |
|
| 94 |
| **IFEval lenient 5Q** | 40% (2/5) | **60% (3/5)** | +20pp |
|
| 95 |
| **IFEval strict 541Q** | ~15% | ~22% | +7pp |
|
| 96 |
+
| **SecEval 2.1k** | 42% | **62.7%** | **+20.7pp** |
|
| 97 |
+
| **HumanEval 10Q** | 12% | **28%** | +16pp |
|
|
|
|
| 98 |
| **CyberMetric 500** | 38% | **67%** | +29pp |
|
| 99 |
+
| **SWE-mini 1 issue** | 0/1 | 0/1 | 0.6B no agéntico (DeepSeek 671B 58.7%) |
|
| 100 |
|
| 101 |
+
*Entreno: 520k líneas (500k rdru200m +20k code) 465MB, 20k steps, r128, DoRA+GaLore+MoE, PT 4x 60→45°C, Levy α1.5, curriculum fácil→difícil.*
|
| 102 |
|
| 103 |
+
### Entrenamiento
|
| 104 |
|
| 105 |
+
Dataset `D:\datasets\rdru200m\parts` 110 shards 6.9GB + `code_search_net` 20k. Ver `src/l8_cyber.c` `cyber_train_particle()`.
|
| 106 |
+
|
| 107 |
+
### Licencia
|
|
|
|
|
|
|
| 108 |
|
| 109 |
+
Apache 2.0 (Qwen3) + gguf2bin MIT
|