AndrewThompson1233 commited on
Commit
b26db35
·
1 Parent(s): 95db57a
Files changed (1) hide show
  1. README.md +30 -12
README.md CHANGED
@@ -9,26 +9,43 @@ pipeline_tag: text-generation
9
 
10
  # Maba v1.5 (103.5M)
11
 
12
- Trained checkpoint built on [AndrewThompson1233/maba-v1.5-exp-architecture](https://huggingface.co/AndrewThompson1233/maba-v1.5-exp-architecture).
 
 
13
 
14
- * Base Architecture: [AndrewThompson1233/maba-v1.5-exp-architecture](https://huggingface.co/AndrewThompson1233/maba-v1.5-exp-architecture)
15
- * Parameters: **103.5M**
16
- * Core Ratio: **95.21%** (4.30% Vocab Tax)
17
- * Macro-Stack: **3:1** (15 DGDA Recurrence : 5 MABA-SA Attention)
18
  * Positional Encoding: **Strict NoPE** (0 parameters)
19
- * Trained on: 3,044 dialogues on NVIDIA L4 (bfloat16)
20
 
21
- ## Empirical Benchmark vs Qwen3.8-Flash-Next (101.7M)
 
 
 
 
22
 
23
- | Metric | Maba v1.5-exp | Qwen3.8-Flash-Next | Delta |
24
  | :--- | :---: | :---: | :---: |
25
- | **Retrieval Accuracy (MCQ)** | **87.5% (7/8)** | 75.0% (6/8) | **+12.5%** |
26
- | **Validation Loss** | **0.0697** | 0.0778 | **-10.4%** |
27
- | **Perplexity (PPL)** | **1.07** | 1.08 | **Maba wins** |
28
- | **Decoding Speed (L4)** | **7.0 tok/s** | 5.5 tok/s | **+27.3% faster** |
 
 
 
 
 
 
 
 
 
29
 
30
  <p align="center"><img src="assets/architecture_comparison.svg" width="920" alt="Benchmark" /></p>
31
 
 
 
32
  ## How to Use
33
 
34
  ```python
@@ -38,6 +55,7 @@ from safetensors.torch import load_file
38
  from transformers import AutoTokenizer
39
 
40
  tokenizer = AutoTokenizer.from_pretrained('gpt2')
 
41
  if 'maba-v1.5-exp-architecture' not in sys.path:
42
  subprocess.run(['git', 'clone', 'https://huggingface.co/AndrewThompson1233/maba-v1.5-exp-architecture'], check=False)
43
  sys.path.insert(0, 'maba-v1.5-exp-architecture')
 
9
 
10
  # Maba v1.5 (103.5M)
11
 
12
+ > [!WARNING]
13
+ > **Research Proof-of-Concept - Not for General / Production Use**
14
+ > This checkpoint is an empirical demonstration and verification artifact. It proves that the **[AndrewThompson1233/maba-v1.5-exp-architecture](https://huggingface.co/AndrewThompson1233/maba-v1.5-exp-architecture)** architecture is fully functional, trainable from scratch, and numerically stable on consumer/enterprise hardware (NVIDIA L4).
15
 
16
+ * Reference Architecture: [AndrewThompson1233/maba-v1.5-exp-architecture](https://huggingface.co/AndrewThompson1233/maba-v1.5-exp-architecture)
17
+ * Parameters: **103,520,911 (103.5M)**
18
+ * Core Computation Ratio: **95.21%** (4.30% Vocab Tax)
19
+ * Topology: **3:1** (15 DGDA Recurrence : 5 MABA-SA Attention)
20
  * Positional Encoding: **Strict NoPE** (0 parameters)
21
+ * Training Setup: 3,044 dialogues, 15 epochs on NVIDIA L4 (bfloat16)
22
 
23
+ ---
24
+
25
+ ## Empirical Benchmark: Maba v1.5 vs Qwen3.8-Flash-Next (~103M)
26
+
27
+ Evaluated under identical training and hardware budgets (3,044 dialogues, bfloat16, NVIDIA L4):
28
 
29
+ | Benchmark / Architecture Metric | 🔵 Maba v1.5-exp | 🟣 Qwen3.8-Flash-Next | Advantage / Delta |
30
  | :--- | :---: | :---: | :---: |
31
+ | **Total Parameters** | **103,520,911 (103.5M)** | 101,701,120 (101.7M) | 0.2% parity |
32
+ | **Core Computation Ratio** | **95.21%** | 74.99% | **+20.2% more active compute** |
33
+ | **Vocab Tax** | **4.30%** (Factorized) | 25.01% (Direct) | **-20.7% parameter bloat** |
34
+ | **Recurrence Engine (75%)** | **DGDA (Decoupled)** | GDN (Standard) | Decoupled erase/write gates |
35
+ | **Attention Engine (25%)** | **MABA-SA (MLA + Top-32)** | QSA (GQA + Micro-block) | 80% KV latent compression |
36
+ | **Positional Encoding** | **Strict NoPE** | 25% Partial RoPE | 0 positional parameters |
37
+ | **Contrastive Retrieval (MCQ)** | **87.5% (7/8)** | 75.0% (6/8) | **+12.5% accuracy** |
38
+ | **Validation Loss** | **0.0697** | 0.0778 | **-10.4% entropy** |
39
+ | **Validation Perplexity (PPL)** | **1.07** | 1.08 | **Maba wins** |
40
+ | **Attention Ablation (PPL Drop)** | **-29.2%** (51.21 -> 36.24) | -5.7% (51.21 -> 48.29) | **MABA-SA cuts error by 29%** |
41
+ | **Needle-in-a-Haystack (4k)** | **91.7% (11/12)** | Lost-in-Middle (0.0) | Hybrid pooling prevents dilution |
42
+ | **KV-Cache @ 4k Context** | **2.50 MB** | 8.00 MB | **-98.8% vs Dense (200 MB)** |
43
+ | **Decode Throughput (L4)** | **7.0 tok/s** | 5.5 tok/s | **+27.3% faster generation** |
44
 
45
  <p align="center"><img src="assets/architecture_comparison.svg" width="920" alt="Benchmark" /></p>
46
 
47
+ ---
48
+
49
  ## How to Use
50
 
51
  ```python
 
55
  from transformers import AutoTokenizer
56
 
57
  tokenizer = AutoTokenizer.from_pretrained('gpt2')
58
+
59
  if 'maba-v1.5-exp-architecture' not in sys.path:
60
  subprocess.run(['git', 'clone', 'https://huggingface.co/AndrewThompson1233/maba-v1.5-exp-architecture'], check=False)
61
  sys.path.insert(0, 'maba-v1.5-exp-architecture')