Instructions to use fla-hub/rwkv7-2.9B-world with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use fla-hub/rwkv7-2.9B-world with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="fla-hub/rwkv7-2.9B-world", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("fla-hub/rwkv7-2.9B-world", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use fla-hub/rwkv7-2.9B-world with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "fla-hub/rwkv7-2.9B-world" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fla-hub/rwkv7-2.9B-world", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/fla-hub/rwkv7-2.9B-world
- SGLang
How to use fla-hub/rwkv7-2.9B-world with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "fla-hub/rwkv7-2.9B-world" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fla-hub/rwkv7-2.9B-world", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "fla-hub/rwkv7-2.9B-world" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "fla-hub/rwkv7-2.9B-world", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use fla-hub/rwkv7-2.9B-world with Docker Model Runner:
docker model run hf.co/fla-hub/rwkv7-2.9B-world
| base_model: | |
| - BlinkDL/rwkv-7-world | |
| language: | |
| - en | |
| - zh | |
| - ja | |
| - ko | |
| - fr | |
| - ar | |
| - es | |
| - pt | |
| license: apache-2.0 | |
| metrics: | |
| - accuracy | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| ```markdown | |
| # rwkv7-2.9B-world | |
| <!-- Provide a quick summary of what the model is/does. --> | |
| This is RWKV-7 model under flash-linear attention format. | |
| ## Model Details | |
| ### Model Description | |
| <!-- Provide a longer summary of what this model is. --> | |
| - **Developed by:** Bo Peng, Yu Zhang, Songlin Yang, Ruichong Zhang | |
| - **Funded by:** RWKV Project (Under LF AI & Data Foundation) | |
| - **Model type:** RWKV7 | |
| - **Language(s) (NLP):** English | |
| - **License:** Apache-2.0 | |
| - **Parameter count:** 2.9B | |
| - **Tokenizer:** RWKV World tokenizer | |
| - **Vocabulary size:** 65,536 | |
| ### Model Sources | |
| <!-- Provide the basic links for the model. --> | |
| - **Repository:** https://github.com/fla-org/flash-linear-attention ; https://github.com/BlinkDL/RWKV-LM | |
| - **Paper:** https://arxiv.org/abs/2503.14456 | |
| ## Uses | |
| <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. --> | |
| Install `flash-linear-attention` and the latest version of `transformers` before using this model: | |
| ```bash | |
| pip install git+https://github.com/fla-org/flash-linear-attention | |
| pip install 'transformers>=4.48.0' | |
| ``` | |
| ### Direct Use | |
| <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. --> | |
| You can use this model just as any other HuggingFace models: | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model = AutoModelForCausalLM.from_pretrained('fla-hub/rwkv7-2.9B-world', trust_remote_code=True) | |
| tokenizer = AutoTokenizer.from_pretrained('fla-hub/rwkv7-2.9B-world', trust_remote_code=True) | |
| model = model.cuda() | |
| prompt = "What is a large language model?" | |
| messages = [ | |
| {"role": "user", "content": "Who are you?"}, | |
| {"role": "assistant", "content": "I am a GPT-3 based model."}, | |
| {"role": "user", "content": prompt} | |
| ] | |
| text = tokenizer.apply_chat_template( | |
| messages, | |
| tokenize=False, | |
| add_generation_prompt=True | |
| ) | |
| model_inputs = tokenizer([text], return_tensors="pt").to(model.device) | |
| generated_ids = model.generate( | |
| **model_inputs, | |
| max_new_tokens=1024, | |
| ) | |
| generated_ids = [ | |
| output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids) | |
| ] | |
| response = tokenizer.batch_decode(generated_ids, skip_special_tokens=False)[0] | |
| print(response) | |
| ``` | |
| ### Training Data | |
| This model is trained on the World v3 with a total of 3.119 trillion tokens. | |
| #### Training Hyperparameters | |
| - **Training regime:** bfloat16, lr 4e-4 to 1e-5 "delayed" cosine decay, wd 0.1 (with increasing batch sizes during the middle) | |
| - **Final Loss:** 1.8745 | |
| - **Token Count:** 3.119 trillion | |
| ## FAQ | |
| Q: safetensors metadata is none. | |
| A: upgrade transformers to >=4.48.0: `pip install 'transformers>=4.48.0'` | |
| ``` |