Text Generation
Transformers
recurrent_qwen
recurrent-depth
latent-reasoning
qwen2.5
research
custom_code
Instructions to use mshapiro123/recurrent-qwen2.5-0.5b-full-block with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mshapiro123/recurrent-qwen2.5-0.5b-full-block with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mshapiro123/recurrent-qwen2.5-0.5b-full-block", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("mshapiro123/recurrent-qwen2.5-0.5b-full-block", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mshapiro123/recurrent-qwen2.5-0.5b-full-block with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mshapiro123/recurrent-qwen2.5-0.5b-full-block" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mshapiro123/recurrent-qwen2.5-0.5b-full-block", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/mshapiro123/recurrent-qwen2.5-0.5b-full-block
- SGLang
How to use mshapiro123/recurrent-qwen2.5-0.5b-full-block with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mshapiro123/recurrent-qwen2.5-0.5b-full-block" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mshapiro123/recurrent-qwen2.5-0.5b-full-block", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mshapiro123/recurrent-qwen2.5-0.5b-full-block" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mshapiro123/recurrent-qwen2.5-0.5b-full-block", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use mshapiro123/recurrent-qwen2.5-0.5b-full-block with Docker Model Runner:
docker model run hf.co/mshapiro123/recurrent-qwen2.5-0.5b-full-block
Download conversion_receipt.json from mshapiro123/recurrent-qwen2.5-0.5b-full-block: direct link, hf CLI and curl.
- Browser
- Download file 1.62 kB
-
https://huggingface.co/mshapiro123/recurrent-qwen2.5-0.5b-full-block/resolve/main/conversion_receipt.json
- Command line
-
hf download hf://mshapiro123/recurrent-qwen2.5-0.5b-full-block/conversion_receipt.json
-
curl -L -o conversion_receipt.json https://huggingface.co/mshapiro123/recurrent-qwen2.5-0.5b-full-block/resolve/main/conversion_receipt.json
1.62 kB
| { | |
| "schema_version": 1, | |
| "kind": "paper_one_hf_checkpoint_conversion", | |
| "status": "green", | |
| "created_at_utc": "2026-08-16T04:38:36.251947Z", | |
| "repo_name": "recurrent-qwen2.5-0.5b-full-block", | |
| "checkpoint_kind": "full_block_delta", | |
| "source_checkpoint": "/content/drive/MyDrive/recurrent-qwen-svgd-checkpoints/stage5_n24_support12_rung_20260707_140139/anneal_to_outcome_final/unfrozen_recurrent_step_6000.pt", | |
| "source_checkpoint_sha256_expected": "898a259db2ab344ece4545e2910b051840e8408dbe4927f799e7cdb3cdd8c7dc", | |
| "source_checkpoint_sha256": "898a259db2ab344ece4545e2910b051840e8408dbe4927f799e7cdb3cdd8c7dc", | |
| "source_phase": "unfrozen_recurrent", | |
| "source_step": 6000, | |
| "source_tensor_count": 152, | |
| "source_total_parameters": 182163457, | |
| "excluded_checkpoint_tensors": { | |
| "bridge.proj.bias": { | |
| "shape": [ | |
| 896 | |
| ], | |
| "parameters": 896 | |
| }, | |
| "bridge.proj.weight": { | |
| "shape": [ | |
| 896, | |
| 1792 | |
| ], | |
| "parameters": 1605632 | |
| } | |
| }, | |
| "excluded_checkpoint_parameters": 1606528, | |
| "exclusion_reason": "Receipt-bound legacy concat projection bypassed by split-mode forward execution", | |
| "safetensors_file": "recurrent_delta.safetensors", | |
| "safetensors_sha256": "095e5466aaf194eefbbfa8fc7c4d63879502d410ba2b03d7a884ab198e2b3831", | |
| "tensor_count": 150, | |
| "total_parameters": 180556929, | |
| "lora_parameters": 0, | |
| "bridge_parameters": 1608321, | |
| "dtype_tensor_counts": { | |
| "bfloat16": 144, | |
| "float32": 6 | |
| }, | |
| "key_transform": "exclude manifest-listed inactive compatibility tensors; base_model.* -> backbone.*; all other release keys unchanged" | |
| } | |