--- library_name: transformers pipeline_tag: text-generation base_model: HuggingFaceTB/SmolLM2-135M tags: - smollm2 - flywire - flyffn - sparse-ffn - connectomics - experimental license: apache-2.0 language: - en - de --- # SmolLM2-135M FlyFFN-v2 **SmolLM2-135M FlyFFN-v2** is an experimental, bio-inspired variant of [HuggingFaceTB/SmolLM2-135M](https://huggingface.co/HuggingFaceTB/SmolLM2-135M). It replaces standard dense Feed-Forward Networks (FFNs) with a **progressive sparse FFN routing mechanism** guided by connectomic principles from the *FlyWire* Drosophila brain graph. This model explores whether sparse, bio-routed activations can maintain or improve zero-shot reasoning efficiency at tiny parameter scales. --- ## 💡 Key Architectural Highlights * **Unchanged Attention Core:** Keeps original multi-head self-attention layers intact from SmolLM2-135M. * **Progressive Sparse FFN Conversion:** FFN layers are initialized from the pretrained dense weights and dynamically sparsified via progressive routing ($8 \rightarrow 6 \rightarrow 4 \rightarrow 3 \rightarrow 2$ shards). * **Dense Anchors:** Strategic layers (`[3, 7, 11, 15, 19, 23, 27]`) remain fully dense to preserve global representation stability. * **Biological Routing Dynamics:** Layer-dependent routing mix and top-$k$ sparsity levels: * **Early/Late Layers:** $k=4$, `mix` $= 0.25 - 0.50$ * **Middle Layers:** $k=6$, `mix` $= 0.10$ --- ## 📊 Evaluation & Performance ### 1. FastEval Zero-Shot Benchmark (50 Samples / Task) *Evaluated on NVIDIA L4 GPU (`torch.bfloat16`). FastEval uses 50 deterministic zero-shot multiple-choice examples per dataset.* | Benchmark | Task Type | SmolLM2-135M (Base) | FlyFFN-v2 | $\Delta$ Difference | | :--- | :--- | :---: | :---: | :---: | | **MMLU-Pro** | Multi-domain Reasoning | 8.0% (4/50) | **10.0%** (5/50) | +2.0% | | **PIQA** | Physical Commonsense | 42.0% (21/50) | **50.0%** (25/50) | +8.0% | | **MMMLU-DE** | Multilingual (German) | 24.0% (12/50) | **32.0%** (16/50) | +8.0% | | **GPQA-Diamond** | High-level Science | *Gated* | *Gated* | — | | **Macro Average** | **Overall** | **24.7%** | **30.7%** | **+6.0%** | > **Note on Throughput:** While sparse FFN routing improves zero-shot accuracy across small-batch evaluations, token generation speed varies by sequence length due to custom routing overheads on standard CUDA kernels (e.g., ~98–110 q/s on PIQA vs. 173 q/s for standard dense). --- ### 2. Language Modeling Metrics | Metric | SmolLM2-135M (Base) | FlyFFN-v2 | | :--- | :---: | :---: | | **Cross-Entropy Loss (CE)** | 2.7227 | 2.9609 | | **Perplexity (PPL)** | 15.22 | 19.32 | --- ## 🛠️ Usage & Reproduction ### Prerequisites & Setup Ensure you have the custom layer definition script `smollm2_flyffn_v2.py` and checkpoint weights `biological_flyffn_v2.pt` in your working directory. ```python import torch from smollm2_flyffn_v2 import SmolLM2FlyFFNv2ForCausalLM from transformers import AutoTokenizer model_id = "HuggingFaceTB/SmolLM2-135M" tokenizer = AutoTokenizer.from_pretrained(model_id) # Load model architecture with custom FlyFFN layers model = SmolLM2FlyFFNv2ForCausalLM.from_pretrained( "./", torch_dtype=torch.bfloat16, device_map="auto" ) input_text = "The physical principle behind gravity is" inputs = tokenizer(input_text, return_tensors="pt").to("cuda") outputs = model.generate(**inputs, max_new_tokens=50) print(tokenizer.decode(outputs[0], skip_special_tokens=True)) ``` ## 🔗 References & Source Code * **Source Repository:** [`vtavakkoli/TinyCeNN-LM`](https://github.com/vtavakkoli/TinyCeNN-LM) * **Base Model:** [`HuggingFaceTB/SmolLM2-135M`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M) * **Evaluation Artifacts:** `fast_eval_50.csv`, `fast_eval_50.json` * **Original Checkpoint:** `biological_flyffn_v2.pt`