apollonlabsai commited on
Commit
00404a6
·
verified ·
1 Parent(s): 1dbb94c

Explain the name Heliactis

Browse files
Files changed (1) hide show
  1. README.md +149 -145
README.md CHANGED
@@ -1,145 +1,149 @@
1
- ---
2
- license: apache-2.0
3
- base_model: XHToken/Spark-X2.5-4B
4
- language:
5
- - el
6
- - en
7
- library_name: llama.cpp
8
- pipeline_tag: text-generation
9
- tags:
10
- - gguf
11
- - greek
12
- - quantized
13
- - tool-calling
14
- ---
15
-
16
- <p align="center">
17
- <img src="https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF/resolve/main/insignia.svg" width="280" alt="ApollonLabsAI — Towards the Sun">
18
- </p>
19
-
20
- # Heliactis-1-4B — GGUF
21
-
22
- Quantized builds of **Heliactis 1**, an Apollon Labs fine-tune of
23
- [`XHToken/Spark-X2.5-4B`](https://huggingface.co/XHToken/Spark-X2.5-4B),
24
- for `llama.cpp`, LM Studio, Ollama and anything else that reads GGUF.
25
-
26
- **What it is for:** Greek that does not fall apart, code that stays intact, and
27
- **tool calling that knows when to stay quiet.**
28
-
29
- ## Files
30
-
31
- | File | Size | Gate |
32
- |---|---|---|
33
- | `Heliactis-1-4B-F16.gguf` | 7.66 GiB | passed |
34
- | `Heliactis-1-4B-Q8_0.gguf` | 4.07 GiB | passed |
35
- | `Heliactis-1-4B-Q6_K.gguf` | 3.15 GiB | passed (imatrix) |
36
- | `Heliactis-1-4B-Q5_K_M.gguf` | 2.77 GiB | passed (imatrix) |
37
- | `Heliactis-1-4B-Q4_K_M.gguf` | 2.42 GiB | passed (imatrix) · **the default**, runs on 6 GB VRAM or plain CPU |
38
-
39
- Every file listed here went through the full gate below. Nothing is shipped unmeasured.
40
-
41
- ## Usage
42
-
43
- ```bash
44
- llama-server -m Heliactis-1-4B-Q4_K_M.gguf --jinja -c 8192 \
45
- --cache-type-k q8_0 --cache-type-v q8_0
46
- ```
47
-
48
- **Context window: 1,048,576 tokens**, inherited unchanged from the base
49
- (`max_position_embeddings`; 27 of 36 layers use a 512-token sliding window, which keeps
50
- the KV cache small). `-c 8192` above is only a default that fits small GPUs: raise it as
51
- far as your memory allows. The KV cache at `q8_0` costs 14,976 bytes per token, about
52
- 14.6 GiB for the full window. Long-context recall was **not re-measured** for this release;
53
- see the limitations below.
54
-
55
- The chat template is embedded in the GGUF. The model was trained and evaluated
56
- with **thinking disabled**. Call it that way. `--jinja` is required for tool
57
- calling, otherwise the toolbox never reaches the model.
58
-
59
- ### Identity
60
-
61
- This model was **not trained on any self-identity data**. If identity matters,
62
- set it in the system prompt:
63
-
64
- ```
65
- You are Heliactis 1, a Greek/English assistant built by Apollon Labs
66
- on top of Spark-X2.5-4B.
67
- ```
68
-
69
- ## A base-model defect, reduced
70
-
71
- On `Spark-X2.5-4B`, 502 vocabulary entries begin with an orphan UTF-8 byte. When the model
72
- emits one, the output breaks and `llama-server` answers HTTP 500, even on an ordinary
73
- translation prompt. We [reported it upstream](https://huggingface.co/XHToken/Spark-X2.5-4B/discussions/25).
74
- Heliactis 1 was trained against those tokens: at the worst position we measured, their
75
- probability fell from **12.7%** (our unreleased intermediate fine-tune, the starting point of
76
- this training) to **0.17%**. Rarer, not impossible.
77
-
78
- **Block them at inference** to close this route (`ban-ids.json` ships in this repo):
79
-
80
- ```bash
81
- llama-server -m Heliactis-1-4B-Q4_K_M.gguf --jinja \
82
- $(python -c "import json;print(' '.join(f'-l {i}-inf' for i in json.load(open('ban-ids.json'))))")
83
- ```
84
-
85
- Via the API, per request: `"logit_bias": [[<id>, false], ...]` for the same ids.
86
- A system prompt does not fix it. Full reproduction, training trade-off and a second
87
- route that blocking does not close: [DEFECT.md](DEFECT.md).
88
-
89
- ## Measured results
90
-
91
- Instrument: 12-part gate, thinking disabled, `--jinja`, KV `q8_0`, `-c 8192`.
92
- Each value comes from the gate's own results files.
93
-
94
- | Test | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | F16 | bar |
95
- |---|---|---|---|---|---|---|
96
- | Tool abstention | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | ≥ 9 |
97
- | Overall tool score | 45/45 | 45/45 | 45/45 | 45/45 | 45/45 | ≥ 28 |
98
- | Tool calls, Greek | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | ≥ 7 |
99
- | Tool calls, English | 20/20 | 20/20 | 20/20 | 20/20 | 20/20 | ≥ 17.5 |
100
- | HumanEval pass@1 | 0.7439 | 0.7622 | 0.7805 | 0.7622 | 0.7866 | ≥ 0.6952 |
101
- | Code fencing | 60/60 | 57/60 | 58/60 | 59/60 | 58/60 | ≥ 53 |
102
- | Capability eval | 52/60 | 53/60 | 55/60 | 54/60 | 54/60 | ≥ 51 |
103
- | Greek character validity | 0.9996 | 0.9983 | 0.9974 | 0.9974 | 0.9968 | ≥ 0.9943 |
104
- | Greek vocabulary | 0.9553 | 0.9536 | 0.9528 | 0.9549 | 0.9568 | ≥ 0.949 |
105
- | Greek meaning | 0.76 | 0.79 | 0.78 | 0.79 | 0.80 | ≥ 0.70 |
106
- | Median answer length (tokens) | 242 | 251.5 | 256.5 | 253.5 | 254.5 | ≤ 266 |
107
- | Truncated answers | 3 | 5 | 2 | 2 | 3 | ≤ 9 |
108
-
109
-
110
- ### Honest limitations
111
-
112
- - **The orphan-byte defect is reduced, not removed.** See above; block the ids to close this route, and see the Lao note below.
113
- - **Do not use this model for Lao.** A second route to the same broken output exists: the model
114
- can emit a token that ends mid-character and then not finish it. Blocking the 502 ids does not
115
- stop this. In our test (greedy, one prompt per language) a Lao paragraph ended in HTTP 500
116
- on this model and on our unreleased intermediate fine-tune, but not on the base model; Thai passed on all three.
117
- - **Lao and Thai got worse.** The 502 blocked tokens are pieces of those scripts.
118
- Perplexity vs our unreleased intermediate fine-tune, fp32, 5 short outside texts (151 tokens,
119
- so the noise is large): English 17.09 → 16.81 · code 18.44 → 19.30 · Greek 21.13 → 27.13 ·
120
- mixed languages 56.73 → 67.90 · Lao/Thai 56.49 → 78.85. The Greek line is one sentence;
121
- the gate's Greek vocabulary test is the measure we trust, and it passed.
122
- - **Greek meaning (0.76 on Q4_K_M)** has a wide ±0.08 band. It detects collapse, not small changes.
123
- - **Abstention was measured on N=15 prompts.** Evidence, not proof.
124
- - **Long context was not re-measured for this release.** The architecture and KV cost
125
- are the base model's (14,976 bytes/token at q8_0; 9 of 36 layers full attention).
126
- An unreleased earlier fine-tune on the same base recalled 3/3 two-hop facts at 231k tokens.
127
- We do not carry that number over as a claim for this model.
128
- - **24 of the 554 code training examples use Greek identifiers.** Known, not yet fixed.
129
-
130
- ## Training
131
-
132
- LoRA (r=16, alpha=32). Stage 1: SFT on 2014 examples (1212 Greek, 554 code,
133
- 248 tool-calling; 742 of the 866 toolbox rows are decoys that teach abstention).
134
- Dataset sha256: `cd58f73ad7fbb339de87a3bef22af6c2ff996b971f829978c78b4f70703f26e6`.
135
- Stage 2: on-policy unlikelihood on the 502 tokens (856 of the model's own answers,
136
- weight 10), with the same SFT set alongside to hold the rest in place.
137
- imatrix: 100 × 512-token chunks of Greek text.
138
-
139
- ## License and attribution
140
-
141
- Apache 2.0, inherited from the base model `XHToken/Spark-X2.5-4B`.
142
- Derivative work by Apollon Labs.
143
-
144
- ---
145
- *Apollon Labs — Towards the Sun.*
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: XHToken/Spark-X2.5-4B
4
+ language:
5
+ - el
6
+ - en
7
+ library_name: llama.cpp
8
+ pipeline_tag: text-generation
9
+ tags:
10
+ - gguf
11
+ - greek
12
+ - quantized
13
+ - tool-calling
14
+ ---
15
+
16
+ <p align="center">
17
+ <img src="https://huggingface.co/ApollonLabs/Heliactis-1-4B-GGUF/resolve/main/insignia.svg" width="280" alt="ApollonLabsAI — Towards the Sun">
18
+ </p>
19
+
20
+ # Heliactis-1-4B — GGUF
21
+
22
+ > **Heliactis** (Greek *ἡλιακτίς*, today *ηλιακτίδα*) means **"sunbeam"**: a name we formed from
23
+ > *hḗlios* ("sun") and *aktís* ("ray"). The Heliactis models are the rays; the larger **Helios**
24
+ > models are the sun itself. The lab is named after **Apollo** (*Apollon*), the Greek god of light.
25
+
26
+ Quantized builds of **Heliactis 1**, an Apollon Labs fine-tune of
27
+ [`XHToken/Spark-X2.5-4B`](https://huggingface.co/XHToken/Spark-X2.5-4B),
28
+ for `llama.cpp`, LM Studio, Ollama and anything else that reads GGUF.
29
+
30
+ **What it is for:** Greek that does not fall apart, code that stays intact, and
31
+ **tool calling that knows when to stay quiet.**
32
+
33
+ ## Files
34
+
35
+ | File | Size | Gate |
36
+ |---|---|---|
37
+ | `Heliactis-1-4B-F16.gguf` | 7.66 GiB | passed |
38
+ | `Heliactis-1-4B-Q8_0.gguf` | 4.07 GiB | passed |
39
+ | `Heliactis-1-4B-Q6_K.gguf` | 3.15 GiB | passed (imatrix) |
40
+ | `Heliactis-1-4B-Q5_K_M.gguf` | 2.77 GiB | passed (imatrix) |
41
+ | `Heliactis-1-4B-Q4_K_M.gguf` | 2.42 GiB | passed (imatrix) · **the default**, runs on 6 GB VRAM or plain CPU |
42
+
43
+ Every file listed here went through the full gate below. Nothing is shipped unmeasured.
44
+
45
+ ## Usage
46
+
47
+ ```bash
48
+ llama-server -m Heliactis-1-4B-Q4_K_M.gguf --jinja -c 8192 \
49
+ --cache-type-k q8_0 --cache-type-v q8_0
50
+ ```
51
+
52
+ **Context window: 1,048,576 tokens**, inherited unchanged from the base
53
+ (`max_position_embeddings`; 27 of 36 layers use a 512-token sliding window, which keeps
54
+ the KV cache small). `-c 8192` above is only a default that fits small GPUs: raise it as
55
+ far as your memory allows. The KV cache at `q8_0` costs 14,976 bytes per token, about
56
+ 14.6 GiB for the full window. Long-context recall was **not re-measured** for this release;
57
+ see the limitations below.
58
+
59
+ The chat template is embedded in the GGUF. The model was trained and evaluated
60
+ with **thinking disabled**. Call it that way. `--jinja` is required for tool
61
+ calling, otherwise the toolbox never reaches the model.
62
+
63
+ ### Identity
64
+
65
+ This model was **not trained on any self-identity data**. If identity matters,
66
+ set it in the system prompt:
67
+
68
+ ```
69
+ You are Heliactis 1, a Greek/English assistant built by Apollon Labs
70
+ on top of Spark-X2.5-4B.
71
+ ```
72
+
73
+ ## A base-model defect, reduced
74
+
75
+ On `Spark-X2.5-4B`, 502 vocabulary entries begin with an orphan UTF-8 byte. When the model
76
+ emits one, the output breaks and `llama-server` answers HTTP 500, even on an ordinary
77
+ translation prompt. We [reported it upstream](https://huggingface.co/XHToken/Spark-X2.5-4B/discussions/25).
78
+ Heliactis 1 was trained against those tokens: at the worst position we measured, their
79
+ probability fell from **12.7%** (our unreleased intermediate fine-tune, the starting point of
80
+ this training) to **0.17%**. Rarer, not impossible.
81
+
82
+ **Block them at inference** to close this route (`ban-ids.json` ships in this repo):
83
+
84
+ ```bash
85
+ llama-server -m Heliactis-1-4B-Q4_K_M.gguf --jinja \
86
+ $(python -c "import json;print(' '.join(f'-l {i}-inf' for i in json.load(open('ban-ids.json'))))")
87
+ ```
88
+
89
+ Via the API, per request: `"logit_bias": [[<id>, false], ...]` for the same ids.
90
+ A system prompt does not fix it. Full reproduction, training trade-off and a second
91
+ route that blocking does not close: [DEFECT.md](DEFECT.md).
92
+
93
+ ## Measured results
94
+
95
+ Instrument: 12-part gate, thinking disabled, `--jinja`, KV `q8_0`, `-c 8192`.
96
+ Each value comes from the gate's own results files.
97
+
98
+ | Test | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | F16 | bar |
99
+ |---|---|---|---|---|---|---|
100
+ | Tool abstention | 15/15 | 15/15 | 15/15 | 15/15 | 15/15 | ≥ 9 |
101
+ | Overall tool score | 45/45 | 45/45 | 45/45 | 45/45 | 45/45 | ≥ 28 |
102
+ | Tool calls, Greek | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | ≥ 7 |
103
+ | Tool calls, English | 20/20 | 20/20 | 20/20 | 20/20 | 20/20 | ≥ 17.5 |
104
+ | HumanEval pass@1 | 0.7439 | 0.7622 | 0.7805 | 0.7622 | 0.7866 | ≥ 0.6952 |
105
+ | Code fencing | 60/60 | 57/60 | 58/60 | 59/60 | 58/60 | ≥ 53 |
106
+ | Capability eval | 52/60 | 53/60 | 55/60 | 54/60 | 54/60 | ≥ 51 |
107
+ | Greek character validity | 0.9996 | 0.9983 | 0.9974 | 0.9974 | 0.9968 | ≥ 0.9943 |
108
+ | Greek vocabulary | 0.9553 | 0.9536 | 0.9528 | 0.9549 | 0.9568 | ≥ 0.949 |
109
+ | Greek meaning | 0.76 | 0.79 | 0.78 | 0.79 | 0.80 | ≥ 0.70 |
110
+ | Median answer length (tokens) | 242 | 251.5 | 256.5 | 253.5 | 254.5 | ≤ 266 |
111
+ | Truncated answers | 3 | 5 | 2 | 2 | 3 | ≤ 9 |
112
+
113
+
114
+ ### Honest limitations
115
+
116
+ - **The orphan-byte defect is reduced, not removed.** See above; block the ids to close this route, and see the Lao note below.
117
+ - **Do not use this model for Lao.** A second route to the same broken output exists: the model
118
+ can emit a token that ends mid-character and then not finish it. Blocking the 502 ids does not
119
+ stop this. In our test (greedy, one prompt per language) a Lao paragraph ended in HTTP 500
120
+ on this model and on our unreleased intermediate fine-tune, but not on the base model; Thai passed on all three.
121
+ - **Lao and Thai got worse.** The 502 blocked tokens are pieces of those scripts.
122
+ Perplexity vs our unreleased intermediate fine-tune, fp32, 5 short outside texts (151 tokens,
123
+ so the noise is large): English 17.09 → 16.81 · code 18.44 → 19.30 · Greek 21.13 → 27.13 ·
124
+ mixed languages 56.73 → 67.90 · Lao/Thai 56.49 → 78.85. The Greek line is one sentence;
125
+ the gate's Greek vocabulary test is the measure we trust, and it passed.
126
+ - **Greek meaning (0.76 on Q4_K_M)** has a wide ±0.08 band. It detects collapse, not small changes.
127
+ - **Abstention was measured on N=15 prompts.** Evidence, not proof.
128
+ - **Long context was not re-measured for this release.** The architecture and KV cost
129
+ are the base model's (14,976 bytes/token at q8_0; 9 of 36 layers full attention).
130
+ An unreleased earlier fine-tune on the same base recalled 3/3 two-hop facts at 231k tokens.
131
+ We do not carry that number over as a claim for this model.
132
+ - **24 of the 554 code training examples use Greek identifiers.** Known, not yet fixed.
133
+
134
+ ## Training
135
+
136
+ LoRA (r=16, alpha=32). Stage 1: SFT on 2014 examples (1212 Greek, 554 code,
137
+ 248 tool-calling; 742 of the 866 toolbox rows are decoys that teach abstention).
138
+ Dataset sha256: `cd58f73ad7fbb339de87a3bef22af6c2ff996b971f829978c78b4f70703f26e6`.
139
+ Stage 2: on-policy unlikelihood on the 502 tokens (856 of the model's own answers,
140
+ weight 10), with the same SFT set alongside to hold the rest in place.
141
+ imatrix: 100 × 512-token chunks of Greek text.
142
+
143
+ ## License and attribution
144
+
145
+ Apache 2.0, inherited from the base model `XHToken/Spark-X2.5-4B`.
146
+ Derivative work by Apollon Labs.
147
+
148
+ ---
149
+ *Apollon Labs — Towards the Sun.*