--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: text-generation datasets: - HuggingFaceFW/fineweb-edu tags: - causal-lm - language-model - base-model - small-language-model - bananamind - bananamind2 - ternary - int8-embeddings - digit-tokenizer - pytorch - safetensors - custom-code - trust-remote-code - custom-architecture --- # TernaryBananaMind-10M **TernaryBananaMind-10M** is a compact decoder-only base language model from BananaMind. Its inference checkpoint stores projection and output weights as packed ternary values and input embeddings as signed 8-bit values. The intended model ID is `BananaMind/TernaryBananaMind-10M`. The model has **8,428,032 logical weights**, a **4,096-token context window**, and a custom **2,048-token digit-aware byte-level BPE tokenizer**. Its `model.safetensors` file is **2,175,676 bytes** (about **2.07 bits per logical weight**, including scales, float normalization weights, and file overhead). The model name is a size class; the exact logical weight count is 8.43M. ## Model Details | Field | Value | |---|---:| | Architecture | BananaMind 2 style decoder-only Transformer | | Logical weights | 8,428,032 | | Layers | 10 | | Hidden size | 256 | | Intermediate size | 704 | | Attention heads | 4 | | KV heads | 2 | | Head dimension | 64 | | Attention | Grouped-query attention with QK norm | | MLP | SwiGLU | | Position embeddings | RoPE, theta 100,000 | | Normalization | RMSNorm, epsilon 1e-6 | | Vocabulary size | 2,048 | | Context length | 4,096 | | Input embeddings | Separate from output head; signed 8-bit with one scale per row | | Projection and output weights | Ternary (-1, 0, +1), five values per byte | | Activations in linear layers | Quantized to 8 bits during inference | | Checkpoint | `model.safetensors`, 2,175,676 bytes | | HF architecture | `BananaAllForCausalLM` | | HF model type | `bananaall` | We have trained TernaryBananaMind-10M on our BananaAll training framework. See it at https://github.com/BananaMind/BananaAll/ to train your own model simply. The packed weights are decoded to temporary tensors for matrix multiplication. **2.07 bits per weight describes checkpoint storage, not arithmetic precision or peak inference memory.** Packed weights are registered as buffers, so `sum(p.numel() for p in model.parameters())` reports only the 6,656 trainable normalization parameters. Use the logical weight count above for model-size comparisons. ## Tokenizer The custom 2k byte-level BPE tokenizer isolates digits during pre-tokenization. | Special token | ID | |---|---:| | `<\|pad\|>` | 0 | | `<\|bos\|>` | 1 | | `<\|eos\|>` | 2 | | `<\|unk\|>` | 3 | ## Training The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAll training framework**. It was trained on a RTX Pro 6000 by Molab and the script was exported as a notebook. The included `training_args.bin` records the settings below. Its maximum-step value is a configured limit, not a separately verified final step count. | Setting | Recorded value | |---|---:| | Maximum optimizer steps | 61,036 | | Per-device micro batch | 8 sequences | | Gradient accumulation | 2 | | Optimizer | PyTorch fused AdamW | | Betas | 0.9, 0.999 | | Peak learning rate | 0.0018 | | Learning-rate schedule | Cosine | | Warmup steps | 1,831 | | Weight decay | 0.01 | | Gradient clipping | 1.0 | | BF16 training | Enabled | | PyTorch compile | Enabled | | Seed | 37 | Training took 2 hours. ## Evaluation The `lm_eval` and BananaMind Base Bench 1.1 scores below were supplied for the **current packed checkpoint**. Harness version, dtype, and runtime settings can affect results. These scores have been evaluated via our BananaAll framework. ### Standard benchmarks (`lm_eval`) | Benchmark | Acc | Acc norm | Samples | |---|---:|---:|---:| | PIQA | 54.46% | 53.54% | 1,838 | | ARC Easy | 30.98% | 31.57% | 2,376 | | ARC Challenge | 18.09% | 21.08% | 1,172 | | HellaSwag | 26.94% | 27.76% | 10,042 | | **Four-task mean** | — | **33.49%** | — | ### BananaMind Base Bench 1.1 The reported official complete run scored **37.14% accuracy (130/350)**, **35.95% weighted accuracy**, and **904 overall Elo**. | Category | Elo | Accuracy | Weighted accuracy | |---|---:|---:|---:| | Language completion | 963 | 58.00% | 58.28% | | Commonsense | 828 | 32.00% | 33.12% | | World knowledge | 869 | 36.00% | 38.24% | | Context tracking | 832 | 32.00% | 28.34% | | Quantitative | 785 | 20.00% | 18.30% | | Logical reasoning | 1,043 | 46.00% | 42.32% | | Code completion | 1,000 | 36.00% | 37.26% | ### ArithMark 3.0 The [official benchmark script](https://huggingface.co/datasets/AxiomicLabs/Arithmark-3.0/raw/main/bencharithmark-3.py) was run against this local packed checkpoint on all **1,000 examples** with its default batch size and context limit, using CUDA and BF16. The model loaded without missing or unexpected checkpoint keys. | Metric | Score | |---|---:| | Raw continuation accuracy | 31.10% (311/1,000) | | Length-normalized accuracy | 31.20% (312/1,000) | Dataset SHA-256: `bf8ab1a5193d52cdf0e05ff0b3ca226bdfcf416cb6e75562dcbe72e7e4559435`. ### Intelligence Index The **Intelligence Index is 4.90**. Each component is adjusted for its chance floor with `N(score, chance) = 100 × (score - chance) / (100 - chance)`. The calculation uses the length-normalized scores above. Combined ARC is the mean of ARC Easy and ARC Challenge before chance adjustment. | Component | Score used | Chance floor | Adjusted score | Weight | |---|---:|---:|---:|---:| | HellaSwag | 27.76% | 25% | 3.68 | 1.00 | | Combined ARC | 26.325% | 25% | 1.77 | 1.00 | | PIQA | 53.54% | 50% | 7.08 | 1.00 | | ArithMark 3.0 | 31.20% | 25% | 8.27 | 0.65 | `(3.68 + 1.7667 + 7.08 + 0.65 × 8.2667) / 3.65 = 4.9041`, rounded to **4.90**. ## Comparison with BananaMind-2-Nano | Property | TernaryBananaMind-10M | BananaMind-2-Nano | |---|---:|---:| | Logical weights | 8,428,032 | 9,968,128 | | Vocabulary size | 2,048 | 8,192 | | Intermediate size | 704 | 768 | | Input/output embeddings | Separate; input is 8-bit | Tied | | Training tokens seen | 2B FineWeb-Edu | 30B across four datasets | | Shared `lm_eval` benchmark | TernaryBananaMind-10M acc norm | BananaMind-2-Nano acc norm | Difference | |---|---:|---:|---:| | PIQA | 53.54% | 55.98% | -2.44 points | | ARC Easy | 31.57% | 36.20% | -4.63 points | | ARC Challenge | 21.08% | 23.38% | -2.30 points | | HellaSwag | 27.76% | 27.50% | +0.26 points | | **Four-task mean** | **33.49%** | **35.77%** | **-2.28 points** | The two models differ in vocabulary, MLP width, training token count, and weight format, so these scores are a model-level comparison rather than an isolated measure of quantization. | Additional metric | TernaryBananaMind-10M | BananaMind-2-Nano | Difference | |---|---:|---:|---:| | Intelligence Index | **4.90** | **8.01** | -3.11 | | ArithMark 3.0 acc norm | 31.20% | 33.70% | -2.50 points | | BananaMind Base Bench 1.1 Elo | 904 | 917 | -13 | | BananaMind Base Bench 1.1 accuracy | 37.14% (130/350) | 40.86% (143/350) | -3.72 points | | BananaMind Base Bench 1.1 weighted accuracy | 35.95% | 37.52% | -1.57 points | ### BananaMind Base Bench 1.1 category comparison | Category | Ternary Elo | Nano Elo | Ternary accuracy | Nano accuracy | Ternary weighted | Nano weighted | |---|---:|---:|---:|---:|---:|---:| | Language completion | 963 | 1,147 | 58.00% | 80.00% | 58.28% | 80.25% | | Commonsense | 828 | 865 | 32.00% | 38.00% | 33.12% | 37.85% | | World knowledge | 869 | 895 | 36.00% | 46.00% | 38.24% | 41.69% | | Context tracking | 832 | 804 | 32.00% | 26.00% | 28.34% | 25.16% | | Quantitative | 785 | 881 | 20.00% | 32.00% | 18.30% | 28.08% | | Logical reasoning | 1,043 | 1,043 | 46.00% | 46.00% | 42.32% | 42.32% | | Code completion | 1,000 | 827 | 36.00% | 18.00% | 37.26% | 18.15% | ## Usage This model uses custom architecture code. Load it with `trust_remote_code=True`. ```bash pip install -U transformers safetensors torch ``` ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "BananaMind/TernaryBananaMind-10M" device = "cuda" if torch.cuda.is_available() else "cpu" dtype = ( torch.bfloat16 if device == "cuda" and torch.cuda.is_bf16_supported() else torch.float32 ) tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained( model_id, trust_remote_code=True, dtype=dtype, ).to(device).eval() inputs = tokenizer("The color of the sky is", return_tensors="pt").to(device) with torch.inference_mode(): output = model.generate( **inputs, max_new_tokens=96, do_sample=True, temperature=0.7, top_p=0.9, pad_token_id=tokenizer.eos_token_id, eos_token_id=tokenizer.eos_token_id, ) print(tokenizer.decode(output[0], skip_special_tokens=True)) ``` ## License Apache 2.0.