--- library_name: transformers pipeline_tag: text-generation license: apache-2.0 language: - "en" - "zh" - "fr" - "es" - "de" - "pt" - "ru" - "it" - "ja" - "ko" - "vi" - "ar" datasets: - "HuggingFaceFW/fineweb-edu" - "mlfoundations/dclm-baseline-1.0" - "cerebras/SlimPajama-627B" - "EleutherAI/pile" - "bigcode/starcoderdata" - "oscar-corpus/OSCAR-2301" tags: - rwkv - rwkv7 - recurrent - causal-lm - conversational ---
--- ## Model introduction This is an official BlinkDL release of **RWKV-7 Goose** in Hugging Face Transformers format. RWKV-7 is an attention-free recurrent architecture with a constant-size recurrent state and constant inference work per generated token. Training remains parallelizable. This checkpoint is a **base model** pretrained with web, code, synthetic, instruction, chat, and reasoning data. It is suitable for evaluation, post-training, and fine-tuning; the included chat template is a prompt interface, not a claim that the checkpoint is a safety-aligned assistant. The Transformers integration, conversion, release packaging, Fast Tokenizer, and optional TileLang inference implementation are distributed with this release. ## Highlights - **Constant recurrent state:** memory does not grow like an attention KV cache. - **Native Transformers layout:** standard config, sharded safetensors, generation, recurrent cache continuation, training, and LoRA workflows. - **Exact Fast Tokenizer:** self-contained Rust-backed `tokenizer.json`, generated from the canonical RWKV World byte vocabulary during conversion. - **Chat-ready:** `chat_template.jinja` supports system, multi-turn, thinking, and strict model-generated tool-call prompts. - **Optional optimized runtime:** the isolated [`inference/`](inference/) bundle provides PyTorch fallback and TileLang acceleration without changing the standard model root. ## Model overview | Field | Value | | --- | --- | | Repository | `RWKV/RWKV7-1.5B-20260805` | | Architecture class | `Rwkv7ForCausalLM` | | Public size label | `1.5`B | | Source parameters | `1,527,668,736` | | Serialized parameters | `1,527,668,736` | | Synthesized compatibility tensors | `0` | | Layers | `24` | | Hidden / FFN size | `2048` / `8192` | | Heads / head size | `32` / `64` | | Vocabulary | `65536` | | Training context | `16384 tokens` | | Weight dtype | `bfloat16` | | Numerical conversion | `source dtype preserved` | | Metadata profile | `g1i` | | Metadata provenance | `locked-profile` | | Source checkpoint | [`BlinkDL/rwkv7-g1/rwkv7-g1i-1.5b-20260805-ctx16384.pth`](https://huggingface.co/BlinkDL/rwkv7-g1/blob/ede85bf8ab2e59aff7d7ca909fbbc73317866d89/rwkv7-g1i-1.5b-20260805-ctx16384.pth) | | Source SHA-256 | `32ef7b5bf4dc8bde843cf26dfad809a1f527e2e76a9e790e7d406e71bcd785da` | ## Transformers quickstart Native `rwkv7` auto-class registration requires Transformers 5.15 or a current source checkout until that release is available. ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "RWKV/RWKV7-1.5B-20260805" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16) ``` The recurrent cache returned by the model can be passed back for incremental decoding. Use an `attention_mask` for padded batches. ## Chat quickstart ```python import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "RWKV/RWKV7-1.5B-20260805" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, dtype=torch.bfloat16, ).to("cuda") messages = [{"role": "user", "content": "Explain why RWKV uses constant state."}] input_ids = tokenizer.apply_chat_template( messages, tokenize=True, add_generation_prompt=True, thinking=False, return_tensors="pt", ).to(model.device) output = model.generate( input_ids, max_new_tokens=256, do_sample=True, temperature=1.0, top_p=0.5, eos_token_id=0, pad_token_id=0, ) print(tokenizer.decode(output[0, input_ids.shape[1]:], skip_special_tokens=True)) ``` Set `thinking=True` for the RWKV thinking prefix. The intentional generation prefixes are `Assistant: