--- license: cc-by-4.0 language: - en tags: - connectomics - drosophila - connectome - text-generation - tiny-model datasets: - fernandofernandes/fly-connectome-49k - roneneldan/TinyStories pipeline_tag: text-generation --- # Fly Wordbrain — rank 32 **A 2,343,125-parameter language model on a frozen fruit-fly connectome — 22.5× smaller than the model it derives from, and still ahead of it on held-out text.** This is the smaller sibling of [fly-wordbrain-rank64](https://huggingface.co/fernandofernandes/fly-wordbrain-rank64), which is the recommended baseline. Everything about the architecture is identical except the readout rank: 49,393 → **32** → 1,024 instead of → 64 →. Research and receipts: [fly-wordbrain](https://github.com/fernando-neto-ai/fly-wordbrain) ## What halving the readout again costs 200 fresh TinyStories, 45,059 next-token targets, selecting no checkpoint. | Model | Parameters | Audit CE | Audit top-1 | ΔCE vs reference (95% CI) | |---|---:|---:|---:|---| | ngxson reference | 52,756,661 | 3.9882 | 31.38% | — | | rank 64 | 3,956,469 | 3.2802 | 33.37% | −0.708 [−0.738, −0.679] | | **rank 32 (this)** | **2,343,125** | **3.3091** | **32.58%** | **−0.679** [−0.711, −0.647] | Relative to rank 64: **1,613,344 fewer parameters** for **0.029 nats** and **0.79 accuracy points**. On the selection-validation split that gap reads as 1.42 points, and its generated text has more visible failures on generic prompts. Take this one when size matters more than the last point of quality. | Component | Parameters | |---|---:| | Encoder (embedding + 8 input projections) | 482,816 | | Neuron gain, recurrent gain, bias | 148,179 | | Output LayerNorm | 98,786 | | Factorized readout (rank 32) | 1,613,344 | | **Total** | **2,343,125** | **No synaptic weight was modified.** The frozen-buffer digests in `manifest.json` are byte-identical to the reference model's and to rank 64's. Training changed the per-neuron dynamics and the two interfaces, nothing else. Both retained selectors land on update 16,600 here, so `min-ce.safetensors` and `max-accuracy.safetensors` hold the same weights under two different selection receipts. ## Caveats The comparison against the released reference is **not** controlled — its training stories and trainer are unknown. The controlled finding is that constraining the readout rank acts as a regularizer against overfitting 1,000 short stories; it says nothing about whether the fly's wiring is a good prior for language. That would need randomized-graph and zero-edge controls, which have not been run. The connectome is not duplicated here; it lives in [fly-connectome-49k](https://huggingface.co/datasets/fernandofernandes/fly-connectome-49k). Loading instructions and the full step function are on the [rank-64 card](https://huggingface.co/fernandofernandes/fly-wordbrain-rank64). ## Licence Weights **CC BY 4.0** — derived from MaleCNS v1.0 (FlyEM / HHMI Janelia, University of Cambridge, MRC LMB, Google Research) and the [`ngxson/fly-llm-hf`](https://huggingface.co/ngxson/fly-llm-hf) architecture. Code MIT. TinyStories is not redistributed.