Instructions to use vtava/SmolLM2-135M-MemoryFusion-Sequential-R64 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/SmolLM2-135M-MemoryFusion-Sequential-R64 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/SmolLM2-135M-MemoryFusion-Sequential-R64")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vtava/SmolLM2-135M-MemoryFusion-Sequential-R64", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/SmolLM2-135M-MemoryFusion-Sequential-R64 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/SmolLM2-135M-MemoryFusion-Sequential-R64" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/SmolLM2-135M-MemoryFusion-Sequential-R64", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/vtava/SmolLM2-135M-MemoryFusion-Sequential-R64
- SGLang
How to use vtava/SmolLM2-135M-MemoryFusion-Sequential-R64 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/SmolLM2-135M-MemoryFusion-Sequential-R64" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/SmolLM2-135M-MemoryFusion-Sequential-R64", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/SmolLM2-135M-MemoryFusion-Sequential-R64" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/SmolLM2-135M-MemoryFusion-Sequential-R64", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use vtava/SmolLM2-135M-MemoryFusion-Sequential-R64 with Docker Model Runner:
docker model run hf.co/vtava/SmolLM2-135M-MemoryFusion-Sequential-R64
Download load_model.py from vtava/SmolLM2-135M-MemoryFusion-Sequential-R64: direct link, hf CLI and curl.
- Browser
- Download file 1.48 kB
-
https://huggingface.co/vtava/SmolLM2-135M-MemoryFusion-Sequential-R64/resolve/main/load_model.py
- Command line
-
hf download hf://vtava/SmolLM2-135M-MemoryFusion-Sequential-R64/load_model.py
-
curl -L -o load_model.py https://huggingface.co/vtava/SmolLM2-135M-MemoryFusion-Sequential-R64/resolve/main/load_model.py
1.48 kB
| from __future__ import annotations | |
| import json | |
| import torch | |
| from huggingface_hub import hf_hub_download | |
| from safetensors.torch import load_file | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| from tinycenn_lm.smollm2_memory_fusion import SmolMemoryFusionConfig, replace_attention_layers | |
| def load_model(repo_id: str, token=None, device=None): | |
| device=torch.device(device or ('cuda' if torch.cuda.is_available() else 'cpu')) | |
| dtype=torch.bfloat16 if device.type=='cuda' and torch.cuda.is_bf16_supported() else (torch.float16 if device.type=='cuda' else torch.float32) | |
| meta_path=hf_hub_download(repo_id,'tinycenn_model.json',token=token) | |
| with open(meta_path,encoding='utf-8') as f: | |
| meta=json.load(f) | |
| tokenizer=AutoTokenizer.from_pretrained(repo_id,token=token,use_fast=True) | |
| if tokenizer.pad_token_id is None: | |
| tokenizer.pad_token=tokenizer.eos_token | |
| model=AutoModelForCausalLM.from_pretrained(meta['base_model'],dtype=dtype) | |
| cfg=SmolMemoryFusionConfig.from_dict(meta['memory_fusion']) | |
| replace_attention_layers(model,cfg,meta['accepted_layers']) | |
| state_path=hf_hub_download(repo_id,'model.safetensors',token=token) | |
| state=load_file(state_path,device='cpu') | |
| model.load_state_dict(state,strict=True) | |
| model.to(device).eval(); model.requires_grad_(False); model.config.use_cache=False | |
| if hasattr(model,'generation_config'): | |
| model.generation_config.use_cache=False | |
| return model,tokenizer,meta | |