Banaxi-Tech commited on
Commit
b7cc374
·
verified ·
1 Parent(s): 603f5fe

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +11 -1
README.md CHANGED
@@ -54,6 +54,9 @@ The model has **8,428,032 logical weights**, a **4,096-token context window**, a
54
  | HF architecture | `BananaAllForCausalLM` |
55
  | HF model type | `bananaall` |
56
 
 
 
 
57
  The packed weights are decoded to temporary tensors for matrix multiplication. **2.07 bits per weight describes checkpoint storage, not arithmetic precision or peak inference memory.** Packed weights are registered as buffers, so `sum(p.numel() for p in model.parameters())` reports only the 6,656 trainable normalization parameters. Use the logical weight count above for model-size comparisons.
58
 
59
  ## Tokenizer
@@ -69,7 +72,7 @@ The custom 2k byte-level BPE tokenizer isolates digits during pre-tokenization.
69
 
70
  ## Training
71
 
72
- The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAll training framework**. The included `training_args.bin` records the settings below. Its maximum-step value is a configured limit, not a separately verified final step count.
73
 
74
  | Setting | Recorded value |
75
  |---|---:|
@@ -87,10 +90,17 @@ The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAl
87
  | PyTorch compile | Enabled |
88
  | Seed | 37 |
89
 
 
 
 
 
90
  ## Evaluation
91
 
92
  The `lm_eval` and BananaMind Base Bench 1.1 scores below were supplied for the **current packed checkpoint**. Harness version, dtype, and runtime settings can affect results.
93
 
 
 
 
94
  ### Standard benchmarks (`lm_eval`)
95
 
96
  | Benchmark | Acc | Acc norm | Samples |
 
54
  | HF architecture | `BananaAllForCausalLM` |
55
  | HF model type | `bananaall` |
56
 
57
+
58
+ We have trained TernaryBananaMind-10M on our BananaAll training framework. See it at https://github.com/BananaMind/BananaAll/ to train your own model simply.
59
+
60
  The packed weights are decoded to temporary tensors for matrix multiplication. **2.07 bits per weight describes checkpoint storage, not arithmetic precision or peak inference memory.** Packed weights are registered as buffers, so `sum(p.numel() for p in model.parameters())` reports only the 6,656 trainable normalization parameters. Use the logical weight count above for model-size comparisons.
61
 
62
  ## Tokenizer
 
72
 
73
  ## Training
74
 
75
+ The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAll training framework**. It was trained on a RTX Pro 6000 by Molab and the script was exported as a notebook. The included `training_args.bin` records the settings below. Its maximum-step value is a configured limit, not a separately verified final step count.
76
 
77
  | Setting | Recorded value |
78
  |---|---:|
 
90
  | PyTorch compile | Enabled |
91
  | Seed | 37 |
92
 
93
+
94
+ Training took 2 hours.
95
+
96
+
97
  ## Evaluation
98
 
99
  The `lm_eval` and BananaMind Base Bench 1.1 scores below were supplied for the **current packed checkpoint**. Harness version, dtype, and runtime settings can affect results.
100
 
101
+
102
+ These scores have been evaluated via our BananaAll framework.
103
+
104
  ### Standard benchmarks (`lm_eval`)
105
 
106
  | Benchmark | Acc | Acc norm | Samples |