Text Generation
Transformers
Safetensors
qwen3_5_text
tinycenn
cenn
language-modeling
research
conversational
Instructions to use vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32") model = AutoModelForCausalLM.from_pretrained("vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32
- SGLang
How to use vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32 with Docker Model Runner:
docker model run hf.co/vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32
Download qwen35_verification.json from vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32: direct link, hf CLI and curl.
- Browser
- Download file 5.12 kB
-
https://huggingface.co/vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32/resolve/main/qwen35_verification.json
- Command line
-
hf download hf://vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32/qwen35_verification.json
-
curl -L -o qwen35_verification.json https://huggingface.co/vtava/Qwen3.5-0.8B-PDelta3-CLVR-Local32/resolve/main/qwen35_verification.json
5.12 kB
| { | |
| "verified": true, | |
| "base_model": "Qwen/Qwen3.5-0.8B", | |
| "accepted_full_attention_layers": [ | |
| 3, | |
| 7, | |
| 11 | |
| ], | |
| "config": { | |
| "feature_dim": 96, | |
| "local_window": 32, | |
| "chunk_size": 32, | |
| "conv_kernel": 4, | |
| "state_dtype": "fp16", | |
| "variant": "conv4_gdn2_clvr_f96", | |
| "local_gate_init": 0.72, | |
| "warm_start_previous_core": true | |
| }, | |
| "probe_context": 128, | |
| "probe_blocks": 6, | |
| "baseline_probe_nll": 2.861148993174235, | |
| "candidate_probe_nll": 2.8818757136662803, | |
| "delta_nll": 0.02072672049204538, | |
| "release_cumulative_delta_nll_limit": 0.05, | |
| "quality_pass": true, | |
| "all_saved_layer_gates_pass": true, | |
| "layer_checks": [ | |
| { | |
| "layer": 3, | |
| "pass": true, | |
| "nmse": 0.062080949544906616, | |
| "cosine": 0.9744633436203003, | |
| "incremental_delta_nll": -0.003246148427327178, | |
| "cumulative_delta_nll": -0.003246148427327178 | |
| }, | |
| { | |
| "layer": 7, | |
| "pass": true, | |
| "nmse": 0.03889689967036247, | |
| "cosine": 0.9704566597938538, | |
| "incremental_delta_nll": 0.013274192810058594, | |
| "cumulative_delta_nll": 0.010028044382731416 | |
| }, | |
| { | |
| "layer": 11, | |
| "pass": true, | |
| "nmse": 0.10417895764112473, | |
| "cosine": 0.9411033391952515, | |
| "incremental_delta_nll": 0.010698676109313965, | |
| "cumulative_delta_nll": 0.02072672049204538 | |
| } | |
| ], | |
| "prompt_examples": [ | |
| { | |
| "prompt": "The future of small language models is", | |
| "baseline": "The future of small language models is not just about the technology, but also about the human side of the conversation.\n\nIn the last few years, small language models (SLMs) have become a game changer in the world of AI. They are fast, cheap, and capable of handling complex tasks. However, they are also not without their limitations.", | |
| "pdelta3_clvr": "The future of small language models is not just about the technology, but also about the human side of the conversation.\n\nIn the last few years, small language models (SLMs) have been gaining traction in the tech industry. They are becoming increasingly popular for their ability to generate text, code, and other tasks. However, they are also facing challenges" | |
| }, | |
| { | |
| "prompt": "Artificial intelligence can help scientists by", | |
| "baseline": "Artificial intelligence can help scientists by providing a way to predict the future.\n\n<think>\n\n</think>\n\nYes, that is a correct statement. Artificial intelligence (AI) offers scientists a powerful tool for **predictive modeling**, which allows them to analyze vast amounts of historical data to identify patterns and trends that would otherwise be difficult to spot.\n\nHere are a", | |
| "pdelta3_clvr": "Artificial intelligence can help scientists by providing a way to predict the future of a system. This is a common example of what?\n\n<think>\n\n</think>\n\nThis is a common example of **artificial intelligence (AI)**.\n\nSpecifically, this scenario describes **machine learning** or **predictive modeling**. In this context, the AI system analyzes historical data" | |
| }, | |
| { | |
| "prompt": "A good software architecture should", | |
| "baseline": "A good software architecture should be able to handle the following:\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.\nA.", | |
| "pdelta3_clvr": "A good software architecture should be able to handle a large number of concurrent users.\nA. True\nB. False\n\n<think>\n\n</think>\n\nThe statement is **True**.\n\nIn modern software development, especially for web applications, mobile devices, and distributed systems, the ability to handle a large number of concurrent users is a fundamental requirement. This" | |
| }, | |
| { | |
| "prompt": "The capital of Austria is", | |
| "baseline": "The capital of Austria is Vienna.\nThe following is a list of the most recent changes to the following:\nThe following is a list of the most recent changes to the following:\nThe following is a list of the most recent changes to the following:\nThe following is a list of the most recent changes to the following:\nThe", | |
| "pdelta3_clvr": "The capital of Austria is Vienna.\nThe following are the most common questions and answers about the topic of \"The Great Wall of China\".\nThe following are the most common questions and answers about the topic of \"The Great Wall of China\".\nThe following are the most common questions and answers about the topic of \"The Great Wall of China" | |
| }, | |
| { | |
| "prompt": "Once upon a time, a small robot", | |
| "baseline": "Once upon a time, a small robot named \"Blinky\" was living in a small room. One day, Blinky decided to play a game with his friends.\nBlinky's friends were:\n- A robot named \"Blinky\"\n- A robot named \"Blinky\"\n- A robot named \"Blinky\"\n- A robot", | |
| "pdelta3_clvr": "Once upon a time, a small robot named \"Blinky\" was exploring the world. One day, he found a mysterious box that contained a special tool. This tool was called a \"safety net\" and it had a special feature.\n\nThe safety net had a special shape. It was like a triangle with a base of 100 units" | |
| } | |
| ] | |
| } |