mlboydaisuke commited on
Commit
4a48b53
·
verified ·
1 Parent(s): 0e8073d

Chat template: accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical

Browse files
README.md CHANGED
@@ -63,6 +63,8 @@ Two ways to read the answer:
63
  | `Shieldstral-1.0-3B-vision_int4.litertlm` | **text + image** | int4-b32 decoder + int8 pixtral tower, static 560×560 | 2.78 GB | 1.82 GiB | **image moderation, phones and up** |
64
  | `Shieldstral-1.0-3B-vision_int4_gpu.litertlm` | **text + image** | the same weights and the same tower; decoder re-exported so the attention softmax lowers to a builtin | 2.78 GB | 1.82 GiB | **the vision file for GPU** |
65
 
 
 
66
  **Pick int4 unless you have a reason not to.** On the gate set it matches int8 on every metric, and it is the only text variant that fits an iPhone: int8's 3.33 GiB single section exceeds the practical iOS mmap budget (~2.1 GiB for an app with default entitlements; even entitlement-relaxed apps have topped out below 3 GiB on current hardware). All variants share an identical int8 embedding section, so the difference is in the decoder weights.
67
 
68
  The vision bundle accepts text documents too, so it can replace the text one — it just costs 0.6 GB more on disk and loads the tower you may not use. **On GPU, use `_int4_gpu` for that.** The original vision file does not load on a mobile GPU at all: litert-torch 0.9.2 marks the attention softmax as an `odml.softmax` StableHLO composite and litert-converter 0.3.0 cannot lower it, so the delegate takes 52 of 1187 ops and the engine is refused. The text files were exported through a path that stripped the marker, which is why they were unaffected. `_int4_gpu` is the same weights and the same pixtral tower, re-exported on litert-converter 0.3.1, which lowers the composite to a builtin `SOFTMAX`.
 
63
  | `Shieldstral-1.0-3B-vision_int4.litertlm` | **text + image** | int4-b32 decoder + int8 pixtral tower, static 560×560 | 2.78 GB | 1.82 GiB | **image moderation, phones and up** |
64
  | `Shieldstral-1.0-3B-vision_int4_gpu.litertlm` | **text + image** | the same weights and the same tower; decoder re-exported so the attention softmax lowers to a builtin | 2.78 GB | 1.82 GiB | **the vision file for GPU** |
65
 
66
+ 2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
67
+
68
  **Pick int4 unless you have a reason not to.** On the gate set it matches int8 on every metric, and it is the only text variant that fits an iPhone: int8's 3.33 GiB single section exceeds the practical iOS mmap budget (~2.1 GiB for an app with default entitlements; even entitlement-relaxed apps have topped out below 3 GiB on current hardware). All variants share an identical int8 embedding section, so the difference is in the decoder weights.
69
 
70
  The vision bundle accepts text documents too, so it can replace the text one — it just costs 0.6 GB more on disk and loads the tower you may not use. **On GPU, use `_int4_gpu` for that.** The original vision file does not load on a mobile GPU at all: litert-torch 0.9.2 marks the attention softmax as an `odml.softmax` StableHLO composite and litert-converter 0.3.0 cannot lower it, so the delegate takes 52 of 1187 ops and the engine is refused. The text files were exported through a path that stripped the marker, which is why they were unaffected. `_int4_gpu` is the same weights and the same pixtral tower, re-exported on litert-converter 0.3.1, which lowers the composite to a builtin `SOFTMAX`.
Shieldstral-1.0-3B-vision_int4.litertlm CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:180b948a4d1d4eacb5219c1124429ee9be392367dcb558e255f7ed466f0f4092
3
  size 2783331824
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3f9289463889fe1232da2bee6ab133804b3e5a381808b00dc1b29e9e36b877bc
3
  size 2783331824
Shieldstral-1.0-3B-vision_int4_gpu.litertlm CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:487639d5e90ff600dd3a9515a9526f11d4f6959c58c2db9bd411c94c1bf6d4d7
3
  size 2783086064
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9dcd46ff64ba528bf57e6153a97c068364efc48c14f36eeb7cada9a56dea7a5b
3
  size 2783086064
litertlm_manifest.json CHANGED
@@ -1,7 +1,7 @@
1
  {
2
  "manifest_schema": "0.1.2",
3
  "repo": "litert-community/Shieldstral-1.0-3B",
4
- "generated": "2026-09-06",
5
  "generator": "make_manifest.py",
6
  "model": {
7
  "display_name": "Shieldstral-1.0-3B",
@@ -21,12 +21,12 @@
21
  "variants": [
22
  {
23
  "file": "Shieldstral-1.0-3B-vision_int4.litertlm",
24
- "sha256": "180b948a4d1d4eacb5219c1124429ee9be392367dcb558e255f7ed466f0f4092",
25
  "size_bytes": 2783331824,
26
  "sections": [
27
  {
28
  "type": "LlmMetadataProto",
29
- "size_bytes": 892
30
  },
31
  {
32
  "type": "HF_Tokenizer_Zlib",
@@ -82,12 +82,12 @@
82
  },
83
  {
84
  "file": "Shieldstral-1.0-3B-vision_int4_gpu.litertlm",
85
- "sha256": "487639d5e90ff600dd3a9515a9526f11d4f6959c58c2db9bd411c94c1bf6d4d7",
86
  "size_bytes": 2783086064,
87
  "sections": [
88
  {
89
  "type": "LlmMetadataProto",
90
- "size_bytes": 892
91
  },
92
  {
93
  "type": "HF_Tokenizer_Zlib",
 
1
  {
2
  "manifest_schema": "0.1.2",
3
  "repo": "litert-community/Shieldstral-1.0-3B",
4
+ "generated": "2026-09-21",
5
  "generator": "make_manifest.py",
6
  "model": {
7
  "display_name": "Shieldstral-1.0-3B",
 
21
  "variants": [
22
  {
23
  "file": "Shieldstral-1.0-3B-vision_int4.litertlm",
24
+ "sha256": "3f9289463889fe1232da2bee6ab133804b3e5a381808b00dc1b29e9e36b877bc",
25
  "size_bytes": 2783331824,
26
  "sections": [
27
  {
28
  "type": "LlmMetadataProto",
29
+ "size_bytes": 1175
30
  },
31
  {
32
  "type": "HF_Tokenizer_Zlib",
 
82
  },
83
  {
84
  "file": "Shieldstral-1.0-3B-vision_int4_gpu.litertlm",
85
+ "sha256": "9dcd46ff64ba528bf57e6153a97c068364efc48c14f36eeb7cada9a56dea7a5b",
86
  "size_bytes": 2783086064,
87
  "sections": [
88
  {
89
  "type": "LlmMetadataProto",
90
+ "size_bytes": 1175
91
  },
92
  {
93
  "type": "HF_Tokenizer_Zlib",