vtava commited on
Commit
cd17aaf
·
verified ·
1 Parent(s): b2851cb

Upload validated TinyCeNN standalone release

Browse files
README.md CHANGED
@@ -1,115 +1,49 @@
1
  ---
2
  library_name: transformers
3
- license: apache-2.0
4
- base_model: Qwen/Qwen3.5-0.8B
5
  tags:
6
- - qwen3.5
7
  - tinycenn
8
- - recurrent-attention
9
- - efficient-attention
10
- - standalone
11
- - safetensors
12
  ---
13
 
14
  # Qwen3.5-0.8B-CeNN-Integrated-V1-Standalone
15
 
16
- Standalone release of `vtava/Qwen3.5-0.8B-CeNN-Integrated-V1` based on **Qwen3.5-0.8B**.
17
-
18
- ## Why this release is different
19
-
20
- This repository contains the **complete model weights**, tokenizer, Qwen3.5 config,
21
- and the minimal TinyCeNN inference runtime required for this architecture. You do
22
- not need to clone the TinyCeNN-LM repository.
23
-
24
- The release was uploaded only after a fresh Python process reconstructed the
25
- custom architecture, loaded `model.safetensors` with strict key checking, and
26
- produced finite logits with the same validation next token as the source model.
27
 
28
  ## Architecture
29
 
30
- - Variant: `cenn_integrated`
31
- - Custom class: `Qwen35IntegratedAttention`
32
- - Replaced full-attention layers: `[3, 23]`
33
- - Parameters: `753,609,552`
34
- - Custom TinyCeNN tensors: `10`
35
- - Standalone reload: **PASS**
36
- - Probe next token: ` Vienna`
37
-
38
- ## Install
39
-
40
- Until Qwen3.5 support is available in a stable Transformers release used by your
41
- environment:
42
-
43
- ```bash
44
- pip install torch safetensors huggingface_hub
45
- pip install git+https://github.com/huggingface/transformers.git@main
46
- ```
47
-
48
- ## Load
49
 
50
- ```python
51
- from huggingface_hub import snapshot_download
52
- from pathlib import Path
53
- import sys
54
 
55
- path = Path(snapshot_download("vtava/Qwen3.5-0.8B-CeNN-Integrated-V1-Standalone"))
56
- sys.path.insert(0, str(path))
57
- from load_model import load_model
58
- model, tokenizer = load_model(path, device="cuda")
59
- ```
60
-
61
- ## TinyCeNN QuickCheck v1
62
-
63
- This is a small deterministic regression/sanity suite designed to finish quickly.
64
- It is **not** an official MMLU, GPQA, IFEval, MMMLU, or LongBench run and its
65
- percentage must not be compared directly with those benchmark scores.
66
 
67
- **Score: 80.0% (4/5)**
68
 
69
- | Category | Result | Prediction | Expected |
70
- |---|---:|---:|---:|
71
- | Knowledge | PASS | B | B |
72
- | STEM | PASS | C | C |
73
- | Reasoning | FAIL | C | D |
74
- | Multilingual | PASS | C | C |
75
- | Context | PASS | B | B |
76
 
77
- Full machine-readable results are in `quickcheck.json`.
 
 
 
 
78
 
79
- ## Qwen3.5-0.8B reference context
80
 
81
- The values below are reference values supplied for context; this release script
82
- does **not** rerun those benchmark suites.
83
 
84
- | Benchmark | Qwen3.5-0.8B |
85
- |---|---:|
86
- | MMLU-Pro (non-thinking) | 29.7 |
87
- | MMLU-Redux (non-thinking) | 48.5 |
88
- | C-Eval (non-thinking) | 46.4 |
89
- | IFEval (non-thinking) | 52.1 |
90
- | MMMLU (non-thinking) | 34.1 |
91
- | MMLU-Pro (thinking) | 42.3 |
92
- | GPQA (thinking) | 11.9 |
93
- | LongBench v2 (thinking) | 26.1 |
94
- | MMMLU (thinking) | 44.3 |
95
- | Global PIQA (thinking) | 59.4 |
96
- | WMT24++ (thinking) | 27.2 |
97
-
98
- ## Files
99
 
100
- - `model.safetensors` — complete reconstructed model weights
101
- - `config.json` — Qwen3.5 model configuration
102
- - tokenizer files
103
- - `tinycenn_lm/` — minimal inference-only TinyCeNN runtime
104
- - `load_model.py` — standalone loader with strict weight checking
105
- - `standalone_config.json` — custom architecture metadata
106
- - `quickcheck.json` — fast sanity-suite results
107
- - `validation_probe.json` — reload-equivalence probe
108
- - `release_report.json` — release summary
109
 
110
- ## Limitations
111
 
112
- - Research model; not a production inference kernel.
113
- - The bundled runtime is reference PyTorch code, not a fused CUDA kernel.
114
- - QuickCheck is only a regression test, not a publication-grade benchmark.
115
- - GGUF still requires native TinyCeNN operator support in llama.cpp.
 
1
  ---
2
  library_name: transformers
3
+ pipeline_tag: text-generation
 
4
  tags:
 
5
  - tinycenn
6
+ - cenn
7
+ - language-modeling
8
+ - text-generation
9
+ - research
10
  ---
11
 
12
  # Qwen3.5-0.8B-CeNN-Integrated-V1-Standalone
13
 
14
+ Research artifact from **TinyCeNN-LM**. Architecture: `TinyCeNN-LM experiment`.
 
 
 
 
 
 
 
 
 
 
15
 
16
  ## Architecture
17
 
18
+ - Architecture/run type: `TinyCeNN-LM experiment`
19
+ - Base model: `not recorded`
20
+ - Dataset: `Not recorded`
21
+ - Source code: https://github.com/vtavakkoli/TinyCeNN-LM
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22
 
23
+ ## Latest saved results
 
 
 
24
 
25
+ | Metric | Value |
26
+ |---|---:|
27
+ | `feature_dim` | 64 |
 
 
 
 
 
 
 
 
28
 
29
+ The Hugging Face repository keeps timestamped run artifacts under `runs/`. This preserves training reports, configs and run metadata independently of the temporary Colab filesystem.
30
 
31
+ ## Saved experiment files
 
 
 
 
 
 
32
 
33
+ - `config.json`
34
+ - `generation_config.json`
35
+ - `release_report.json`
36
+ - `standalone_config.json`
37
+ - `tokenizer_config.json`
38
 
39
+ ## Reproducibility
40
 
41
+ Run the matching notebook from the TinyCeNN-LM repository. Colab notebooks use a Hugging Face write token from the `HF_TOKEN` Colab Secret; tokens should never be pasted into notebook source.
 
42
 
43
+ ## Limitations
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
+ This is a research checkpoint. Metrics saved here are the metrics produced by the corresponding training notebook/script; unless explicitly marked as held-out evaluation, they should not be treated as publication-grade benchmark results. Generation quality can differ substantially from the base model.
 
 
 
 
 
 
 
 
46
 
47
+ ## Citation
48
 
49
+ If you use this experimental checkpoint, cite the TinyCeNN-LM repository and the upstream base model.
 
 
 
__pycache__/load_model.cpython-313.pyc CHANGED
Binary files a/__pycache__/load_model.cpython-313.pyc and b/__pycache__/load_model.cpython-313.pyc differ
 
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:b601d5794f55ceb4f8cdfc2e621defa235da3226d129fdc18b9583077812765a
3
  size 1509689856
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3d66e5c849ac49974a41444662995ca4fe23c6ef6b07bdcbaeb2fbe74cb5193d
3
  size 1509689856
run_manifest.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "run_id": "vtava__Qwen3.5-0.8B-CeNN-20260917T221342Z",
3
+ "created_utc": "2026-09-17T22:13:42.455764+00:00",
4
+ "repo_id": "vtava/Qwen3.5-0.8B-CeNN-Integrated-V1-Standalone",
5
+ "notebook": null,
6
+ "python": "3.13.15",
7
+ "platform": "Linux-6.6.122+-x86_64-with-glibc2.39",
8
+ "reports": [
9
+ "config.json",
10
+ "generation_config.json",
11
+ "release_report.json",
12
+ "standalone_config.json",
13
+ "tokenizer_config.json"
14
+ ],
15
+ "torch": "2.11.0+cu128",
16
+ "cuda_available": true,
17
+ "gpu": "Tesla T4"
18
+ }
tinycenn_lm/__pycache__/__init__.cpython-313.pyc CHANGED
Binary files a/tinycenn_lm/__pycache__/__init__.cpython-313.pyc and b/tinycenn_lm/__pycache__/__init__.cpython-313.pyc differ
 
tinycenn_lm/__pycache__/optimized_memory.cpython-313.pyc CHANGED
Binary files a/tinycenn_lm/__pycache__/optimized_memory.cpython-313.pyc and b/tinycenn_lm/__pycache__/optimized_memory.cpython-313.pyc differ
 
tinycenn_lm/__pycache__/qwen35_integrated_memory.cpython-313.pyc CHANGED
Binary files a/tinycenn_lm/__pycache__/qwen35_integrated_memory.cpython-313.pyc and b/tinycenn_lm/__pycache__/qwen35_integrated_memory.cpython-313.pyc differ