Text Generation
Transformers
Safetensors
qwen3_5_text
tinycenn
cenn
language-modeling
research
conversational
Instructions to use vtava/Qwen3.5-0.8B-MemoryFusion-Standalone with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/Qwen3.5-0.8B-MemoryFusion-Standalone with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/Qwen3.5-0.8B-MemoryFusion-Standalone") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("vtava/Qwen3.5-0.8B-MemoryFusion-Standalone", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/Qwen3.5-0.8B-MemoryFusion-Standalone with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/Qwen3.5-0.8B-MemoryFusion-Standalone" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Qwen3.5-0.8B-MemoryFusion-Standalone", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vtava/Qwen3.5-0.8B-MemoryFusion-Standalone
- SGLang
How to use vtava/Qwen3.5-0.8B-MemoryFusion-Standalone with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/Qwen3.5-0.8B-MemoryFusion-Standalone" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Qwen3.5-0.8B-MemoryFusion-Standalone", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/Qwen3.5-0.8B-MemoryFusion-Standalone" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Qwen3.5-0.8B-MemoryFusion-Standalone", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vtava/Qwen3.5-0.8B-MemoryFusion-Standalone with Docker Model Runner:
docker model run hf.co/vtava/Qwen3.5-0.8B-MemoryFusion-Standalone
Download fast_eval.json from vtava/Qwen3.5-0.8B-MemoryFusion-Standalone: direct link, hf CLI and curl.
- Browser
- Download file 4.14 kB
-
https://huggingface.co/vtava/Qwen3.5-0.8B-MemoryFusion-Standalone/resolve/main/fast_eval.json
- Command line
-
hf download hf://vtava/Qwen3.5-0.8B-MemoryFusion-Standalone/fast_eval.json
-
curl -L -o fast_eval.json https://huggingface.co/vtava/Qwen3.5-0.8B-MemoryFusion-Standalone/resolve/main/fast_eval.json
4.14 kB
| { | |
| "suite": "TinyCeNN FastEval v1", | |
| "official_full_benchmark": false, | |
| "method": "deterministic sampled zero-shot next-token letter scoring", | |
| "samples_per_benchmark": 50, | |
| "seed": 42, | |
| "datasets": { | |
| "MMLU-Pro": "TIGER-Lab/MMLU-Pro:test", | |
| "PIQA": "lighteval/piqa:validation", | |
| "MMMLU-DE": "openai/MMMLU:DE_DE:test", | |
| "GPQA-Diamond": "Wanfq/gpqa:gpqa_diamond:train" | |
| }, | |
| "device": "cuda", | |
| "torch_version": "2.11.0+cu128", | |
| "results": [ | |
| { | |
| "name": "vtava__Qwen3.5-0.8B-MemoryFusion-Standalone", | |
| "benchmarks": [ | |
| { | |
| "benchmark": "MMLU-Pro", | |
| "samples": 50, | |
| "correct": 5, | |
| "accuracy_pct": 10.0, | |
| "seconds": 16.41, | |
| "items_per_second": 3.046, | |
| "input_tokens": 11865 | |
| }, | |
| { | |
| "benchmark": "PIQA", | |
| "samples": 50, | |
| "correct": 27, | |
| "accuracy_pct": 54.0, | |
| "seconds": 7.66, | |
| "items_per_second": 6.529, | |
| "input_tokens": 3588 | |
| }, | |
| { | |
| "benchmark": "MMMLU-DE", | |
| "samples": 50, | |
| "correct": 14, | |
| "accuracy_pct": 28.0, | |
| "seconds": 8.62, | |
| "items_per_second": 5.802, | |
| "input_tokens": 4562 | |
| }, | |
| { | |
| "benchmark": "GPQA-Diamond", | |
| "samples": 50, | |
| "correct": 15, | |
| "accuracy_pct": 30.0, | |
| "seconds": 15.72, | |
| "items_per_second": 3.18, | |
| "input_tokens": 10910 | |
| } | |
| ], | |
| "overall": { | |
| "samples": 200, | |
| "correct": 61, | |
| "accuracy_pct": 30.5, | |
| "seconds": 48.41, | |
| "input_tokens": 30925 | |
| }, | |
| "generation": { | |
| "prompts": 3, | |
| "generated_tokens": 72, | |
| "seconds": 8.094, | |
| "tokens_per_second": 8.895, | |
| "peak_vram_gib": 1.51, | |
| "sample_outputs": [ | |
| "\n\n<think>\n\n</think>\n\nVienna is the capital of Austria because it serves as the central political, cultural, and administrative", | |
| "\n\nTo determine how long the robot can operate, we need to calculate the total number of hours it can run on one", | |
| "\n\n<think>\n\n</think>\n\nAn **API gateway** is a central server that acts as a bridge between a web application (" | |
| ] | |
| }, | |
| "peak_vram_gib_observed": 1.51 | |
| }, | |
| { | |
| "name": "Qwen/Qwen3.5-0.8B", | |
| "benchmarks": [ | |
| { | |
| "benchmark": "MMLU-Pro", | |
| "samples": 50, | |
| "correct": 6, | |
| "accuracy_pct": 12.0, | |
| "seconds": 10.51, | |
| "items_per_second": 4.756, | |
| "input_tokens": 11865 | |
| }, | |
| { | |
| "benchmark": "PIQA", | |
| "samples": 50, | |
| "correct": 24, | |
| "accuracy_pct": 48.0, | |
| "seconds": 4.68, | |
| "items_per_second": 10.674, | |
| "input_tokens": 3588 | |
| }, | |
| { | |
| "benchmark": "MMMLU-DE", | |
| "samples": 50, | |
| "correct": 17, | |
| "accuracy_pct": 34.0, | |
| "seconds": 5.45, | |
| "items_per_second": 9.17, | |
| "input_tokens": 4562 | |
| }, | |
| { | |
| "benchmark": "GPQA-Diamond", | |
| "samples": 50, | |
| "correct": 13, | |
| "accuracy_pct": 26.0, | |
| "seconds": 9.65, | |
| "items_per_second": 5.181, | |
| "input_tokens": 10910 | |
| } | |
| ], | |
| "overall": { | |
| "samples": 200, | |
| "correct": 60, | |
| "accuracy_pct": 30.0, | |
| "seconds": 30.3, | |
| "input_tokens": 30925 | |
| }, | |
| "generation": { | |
| "prompts": 3, | |
| "generated_tokens": 72, | |
| "seconds": 5.225, | |
| "tokens_per_second": 13.781, | |
| "peak_vram_gib": 1.473, | |
| "sample_outputs": [ | |
| "\n\n<think>\n\n</think>\n\nVienna is the capital of Austria because it is the largest city in the country and serves as", | |
| "\n\nTo determine how long the robot can operate, we need to calculate the total number of hours it can run on one", | |
| "\n\n<think>\n\n</think>\n\nAn **API Gateway** is a central server that acts as a single entry point for all incoming" | |
| ] | |
| }, | |
| "peak_vram_gib_observed": 1.473 | |
| } | |
| ] | |
| } |