Text Generation
Transformers
Safetensors
PyTorch
English
bananaall
causal-lm
language-model
base-model
small-language-model
bananamind
bananamind2
ternary
int8-embeddings
digit-tokenizer
custom-code
trust-remote-code
custom-architecture
custom_code
8-bit precision
Instructions to use BananaMind/TernaryBananaMind-10M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BananaMind/TernaryBananaMind-10M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="BananaMind/TernaryBananaMind-10M", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("BananaMind/TernaryBananaMind-10M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use BananaMind/TernaryBananaMind-10M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "BananaMind/TernaryBananaMind-10M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BananaMind/TernaryBananaMind-10M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/BananaMind/TernaryBananaMind-10M
- SGLang
How to use BananaMind/TernaryBananaMind-10M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "BananaMind/TernaryBananaMind-10M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BananaMind/TernaryBananaMind-10M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "BananaMind/TernaryBananaMind-10M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "BananaMind/TernaryBananaMind-10M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use BananaMind/TernaryBananaMind-10M with Docker Model Runner:
docker model run hf.co/BananaMind/TernaryBananaMind-10M
Update README.md
Browse files
README.md
CHANGED
|
@@ -54,6 +54,9 @@ The model has **8,428,032 logical weights**, a **4,096-token context window**, a
|
|
| 54 |
| HF architecture | `BananaAllForCausalLM` |
|
| 55 |
| HF model type | `bananaall` |
|
| 56 |
|
|
|
|
|
|
|
|
|
|
| 57 |
The packed weights are decoded to temporary tensors for matrix multiplication. **2.07 bits per weight describes checkpoint storage, not arithmetic precision or peak inference memory.** Packed weights are registered as buffers, so `sum(p.numel() for p in model.parameters())` reports only the 6,656 trainable normalization parameters. Use the logical weight count above for model-size comparisons.
|
| 58 |
|
| 59 |
## Tokenizer
|
|
@@ -69,7 +72,7 @@ The custom 2k byte-level BPE tokenizer isolates digits during pre-tokenization.
|
|
| 69 |
|
| 70 |
## Training
|
| 71 |
|
| 72 |
-
The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAll training framework**. The included `training_args.bin` records the settings below. Its maximum-step value is a configured limit, not a separately verified final step count.
|
| 73 |
|
| 74 |
| Setting | Recorded value |
|
| 75 |
|---|---:|
|
|
@@ -87,10 +90,17 @@ The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAl
|
|
| 87 |
| PyTorch compile | Enabled |
|
| 88 |
| Seed | 37 |
|
| 89 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 90 |
## Evaluation
|
| 91 |
|
| 92 |
The `lm_eval` and BananaMind Base Bench 1.1 scores below were supplied for the **current packed checkpoint**. Harness version, dtype, and runtime settings can affect results.
|
| 93 |
|
|
|
|
|
|
|
|
|
|
| 94 |
### Standard benchmarks (`lm_eval`)
|
| 95 |
|
| 96 |
| Benchmark | Acc | Acc norm | Samples |
|
|
|
|
| 54 |
| HF architecture | `BananaAllForCausalLM` |
|
| 55 |
| HF model type | `bananaall` |
|
| 56 |
|
| 57 |
+
|
| 58 |
+
We have trained TernaryBananaMind-10M on our BananaAll training framework. See it at https://github.com/BananaMind/BananaAll/ to train your own model simply.
|
| 59 |
+
|
| 60 |
The packed weights are decoded to temporary tensors for matrix multiplication. **2.07 bits per weight describes checkpoint storage, not arithmetic precision or peak inference memory.** Packed weights are registered as buffers, so `sum(p.numel() for p in model.parameters())` reports only the 6,656 trainable normalization parameters. Use the logical weight count above for model-size comparisons.
|
| 61 |
|
| 62 |
## Tokenizer
|
|
|
|
| 72 |
|
| 73 |
## Training
|
| 74 |
|
| 75 |
+
The model was trained on **2 billion tokens of FineWeb-Edu** with the **BananaAll training framework**. It was trained on a RTX Pro 6000 by Molab and the script was exported as a notebook. The included `training_args.bin` records the settings below. Its maximum-step value is a configured limit, not a separately verified final step count.
|
| 76 |
|
| 77 |
| Setting | Recorded value |
|
| 78 |
|---|---:|
|
|
|
|
| 90 |
| PyTorch compile | Enabled |
|
| 91 |
| Seed | 37 |
|
| 92 |
|
| 93 |
+
|
| 94 |
+
Training took 2 hours.
|
| 95 |
+
|
| 96 |
+
|
| 97 |
## Evaluation
|
| 98 |
|
| 99 |
The `lm_eval` and BananaMind Base Bench 1.1 scores below were supplied for the **current packed checkpoint**. Harness version, dtype, and runtime settings can affect results.
|
| 100 |
|
| 101 |
+
|
| 102 |
+
These scores have been evaluated via our BananaAll framework.
|
| 103 |
+
|
| 104 |
### Standard benchmarks (`lm_eval`)
|
| 105 |
|
| 106 |
| Benchmark | Acc | Acc norm | Samples |
|