Text Generation
Transformers
Safetensors
Russian
English
zarya
feature-extraction
dllm
diffusion
diffusion-language-modeling
instruct
conversational
custom_code
Instructions to use ai-forever/Zarya-1.7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ai-forever/Zarya-1.7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ai-forever/Zarya-1.7B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("ai-forever/Zarya-1.7B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ai-forever/Zarya-1.7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ai-forever/Zarya-1.7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-forever/Zarya-1.7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ai-forever/Zarya-1.7B
- SGLang
How to use ai-forever/Zarya-1.7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ai-forever/Zarya-1.7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-forever/Zarya-1.7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ai-forever/Zarya-1.7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ai-forever/Zarya-1.7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ai-forever/Zarya-1.7B with Docker Model Runner:
docker model run hf.co/ai-forever/Zarya-1.7B
|
Download README.md from ai-forever/Zarya-1.7B: direct link, hf CLI and curl.
- Browser
- Download file 14.4 kB
-
https://huggingface.co/ai-forever/Zarya-1.7B/resolve/main/README.md
- Command line
-
hf download hf://ai-forever/Zarya-1.7B/README.md
-
curl -L -o README.md https://huggingface.co/ai-forever/Zarya-1.7B/resolve/main/README.md
14.4 kB
| license: mit | |
| language: | |
| - ru | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - dllm | |
| - diffusion | |
| - diffusion-language-modeling | |
| - instruct | |
| library_name: transformers | |
| # Zarya-1.7B | |
| Zarya is a family of hybrid language models that combine a classic auto-regressive (AR) objective with a masked-diffusion (MDM) objective in one model. | |
| The architecture can be built on top of any autoregressive model but in this repository it uses the `Qwen3` backbone. | |
| Naming explanation: Zarya (pronounced as [zɐˈrʲa] ([IPA notation](https://en.wiktionary.org/wiki/Appendix:Russian_pronunciation)), literally "Dawn" in English) is a figure from Slavic folklore — a female personification of dawn who may be considered a goddess. | |
| In various traditions, she can manifest as a single being or as two or three sisters simultaneously. | |
| This is a research prototype. | |
| ## Model Details | |
| ### Model Description | |
| Zarya is a research prototype of a family of hybrid language models that jointly learn a classic auto-regressive (AR) objective and a masked-diffusion (MDM) objective within a single model. | |
| Two generation modes are supported, both reachable through a single `model.generate(...)` call: masked-diffusion (MDM) sampling and slotted-level speculative parallel decoding. | |
| - **Model type:** Hybrid auto-regressive (AR) + masked-diffusion language model (DLLM); backbone `Qwen3`, wrapper `Zarya` | |
| - **Language(s) (NLP):** Russian and English | |
| - **License:** MIT | |
| - **Preprint:** https://arxiv.org/abs/2609.19868 | |
| - **Repository with training code:** https://github.com/ai-forever/zarya | |
| ### Zarya-1.7B details | |
| Zarya-1.7B has the following features: | |
| | Variant | hidden_size | num_hidden_layers | num_attention_heads | intermediate_size | | |
| |------------|-------------|-------------------|---------------------|-------------------| | |
| | Zarya-1.7B | 2048 | 28 | 16 | 6144 | | |
| Context Length: 2048 | |
| ## Uses | |
| Zarya is intended for text generation. | |
| It supports conversational fine-tuning (SFT) and classic auto-regressive pretraining. | |
| ### Direct Use | |
| Direct use is text generation (continuation of a prompt) through the `model.generate(...)` interface, including chat-style prompts formatted with the provided chat template. | |
| Two inference modes are available through the same `generate()` call. | |
| Both modes fully use the KV cache with causal attention masks. | |
| - **MDM sampling** (`slotted_generation=false`): iterative masked-diffusion denoising with the first-hitting sampler. | |
| - **Slotted speculative decoding** (`slotted_generation=true`): parallel slot generation with inter-slot diffusion-based selection and intra-slot autoregressive generation for a decoding speedup. | |
| ### Out-of-Scope Use | |
| The model is a research prototype. | |
| It should not be used for production decisions, safety-critical applications, or any use case where accuracy and reliability are essential without additional evaluation and safeguards. | |
| Inference performance and stability also depend on the chosen decoding hyperparameters (like `slotted_generation`, `slot_size`, `serial_num_blocks`, `slot_threshold`, `token_threshold`, and others). | |
| ## Bias, Risks, and Limitations | |
| This is a research prototype. | |
| The code relies on Hugging Face Transformers APIs; when upgrading versions, compatibility must be checked (tested on Transformers 5.12.1 and PyTorch 2.9.0). | |
| Inference performance and stability depend on the choice of config parameters. | |
| ## How to Get Started with the Model | |
| Use the code below to get started with the model. Loading the model and tokenizer requires `trust_remote_code=True`. | |
| ```python | |
| import torch | |
| from transformers import AutoModel, AutoTokenizer | |
| model_name = "ai-forever/Zarya-1.7B" | |
| model = AutoModel.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.bfloat16).cuda() | |
| tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True) | |
| prompt = "<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n" | |
| input_ids = tokenizer(prompt, return_tensors="pt").input_ids.cuda() | |
| # Both modes go through model.generate(...). | |
| # With generation_config.slotted_generation=true -> slotted speculative decoding: | |
| out = model.generate( | |
| input_ids, | |
| max_new_tokens=256, | |
| do_sample=True, | |
| temperature=0.7, | |
| slot_size=16, | |
| serial_num_blocks=4, | |
| slot_threshold=0.9, | |
| token_threshold=0.3, | |
| ) | |
| # Setting generation_config.slotted_generation=false -> MDM sampling instead: | |
| # out = model.generate(input_ids, max_new_tokens=256) | |
| print(tokenizer.decode(out[0])) | |
| ``` | |
| ## Evaluation | |
| LM-eval benchmarking with the `lm-eval` package is supported. Example run: | |
| ```bash | |
| lm_eval run \ | |
| --tasks=gsm8k,ifeval,mbpp,mbpp_instruct,mbpp_plus,mbpp_plus_instruct,hellaswag \ | |
| --model=hf --confirm_run_unsafe_code \ | |
| --log_samples \ | |
| --apply_chat_template \ | |
| --output_path=./reports/lm-eval_results \ | |
| --model_args=pretrained=ai-forever/Zarya-1.7B,backend=causal,dtype=bfloat16,attn_implementation=sdpa,trust_remote_code=True \ | |
| --gen_kwargs slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4 | |
| ``` | |
| ### Zarya-1.7B results | |
| Measurements below were collected with varying inference parameters and on different GPUs; performance is sensitive to both, so results may differ across configurations and hardware setups. | |
| #### A100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4` | |
| Hardware info: | |
| gpu_driver_cuda_version 13.2; | |
| gpu_driver_version 595.71.05; | |
| Docker info: | |
| Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; | |
| `lm-eval == 0.4.12` | |
| | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | | |
| |-----------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | |
| | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.4670 | ± | 0.0137 | | |
| | | | strict-match | 5 | exact_match | ↑ | 0.4685 | ± | 0.0137 | | |
| | hellaswag | 1 | none | 0 | acc | ↑ | 0.4224 | ± | 0.0049 | | |
| | | | none | 0 | acc_norm | ↑ | 0.5472 | ± | 0.0050 | | |
| | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.5815 | ± | N/A | | |
| | | | none | 0 | inst_level_strict_acc | ↑ | 0.5552 | ± | N/A | | |
| | | | none | 0 | prompt_level_loose_acc | ↑ | 0.4492 | ± | 0.0214 | | |
| | | | none | 0 | prompt_level_strict_acc | ↑ | 0.4288 | ± | 0.0213 | | |
| | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.3540 | ± | 0.0214 | | |
| | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.4815 | ± | 0.0257 | | |
| #### H100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4` | |
| Hardware info: | |
| gpu_driver_cuda_version 13.0; | |
| gpu_driver_version 580.105.08; | |
| Docker info: | |
| Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; | |
| `lm-eval == 0.4.12` | |
| | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | | |
| |--------------------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | |
| | arc_challenge | 1 | none | 0 | acc | ↑ | 0.4437 | ± | 0.0145 | | |
| | | | none | 0 | acc_norm | ↑ | 0.4573 | ± | 0.0146 | | |
| | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.0553 | ± | 0.0063 | | |
| | | | strict-match | 5 | exact_match | ↑ | 0.0455 | ± | 0.0057 | | |
| | hellaswag | 1 | none | 0 | acc | ↑ | 0.4227 | ± | 0.0049 | | |
| | | | none | 0 | acc_norm | ↑ | 0.5474 | ± | 0.0050 | | |
| | hendrycks_math500 | 1 | none | 0 | exact_match | ↑ | 0.0040 | ± | 0.0028 | | |
| | humaneval | 1 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | |
| | humaneval_instruct | 4 | create_test | 0 | pass@1 | ↑ | 0.0549 | ± | 0.0178 | | |
| | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.4317 | ± | N/A | | |
| | | | none | 0 | inst_level_strict_acc | ↑ | 0.4053 | ± | N/A | | |
| | | | none | 0 | prompt_level_loose_acc | ↑ | 0.3013 | ± | 0.0197 | | |
| | | | none | 0 | prompt_level_strict_acc | ↑ | 0.2791 | ± | 0.0193 | | |
| | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | |
| | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | |
| #### H100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=false,T=0,temperature=0.5,top_p=0.8,do_sample=true,noise_schedule=linear` | |
| Hardware info: | |
| gpu_driver_cuda_version 13.0; | |
| gpu_driver_version 580.105.08; | |
| Docker info: | |
| Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; | |
| `lm-eval == 0.4.12` | |
| | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | | |
| |--------------------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | |
| | arc_challenge | 1 | none | 0 | acc | ↑ | 0.4437 | ± | 0.0145 | | |
| | | | none | 0 | acc_norm | ↑ | 0.4573 | ± | 0.0146 | | |
| | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.0091 | ± | 0.0026 | | |
| | | | strict-match | 5 | exact_match | ↑ | 0.0000 | ± | 0 | | |
| | hellaswag | 1 | none | 0 | acc | ↑ | 0.4227 | ± | 0.0049 | | |
| | | | none | 0 | acc_norm | ↑ | 0.5474 | ± | 0.0050 | | |
| | hendrycks_math500 | 1 | none | 0 | exact_match | ↑ | 0.0000 | ± | 0 | | |
| | humaneval | 1 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | |
| | humaneval_instruct | 4 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | |
| | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.2002 | ± | N/A | | |
| | | | none | 0 | inst_level_strict_acc | ↑ | 0.1655 | ± | N/A | | |
| | | | none | 0 | prompt_level_loose_acc | ↑ | 0.1072 | ± | 0.0133 | | |
| | | | none | 0 | prompt_level_strict_acc | ↑ | 0.0776 | ± | 0.0115 | | |
| | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | |
| | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | |
| #### H100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4,max_gen_toks=2048` | |
| Hardware info: | |
| gpu_driver_cuda_version 13.0; | |
| gpu_driver_version 580.105.08; | |
| Docker info: | |
| Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; | |
| `lm-eval == 0.4.12` | |
| | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | | |
| |--------------------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | |
| | arc_challenge | 1 | none | 0 | acc | ↑ | 0.4437 | ± | 0.0145 | | |
| | | | none | 0 | acc_norm | ↑ | 0.4573 | ± | 0.0146 | | |
| | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.0167 | ± | 0.0035 | | |
| | | | strict-match | 5 | exact_match | ↑ | 0.0099 | ± | 0.0027 | | |
| | hellaswag | 1 | none | 0 | acc | ↑ | 0.4227 | ± | 0.0049 | | |
| | | | none | 0 | acc_norm | ↑ | 0.5474 | ± | 0.0050 | | |
| | hendrycks_math500 | 1 | none | 0 | exact_match | ↑ | 0.0000 | ± | 0 | | |
| | humaneval | 1 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | |
| | humaneval_instruct | 4 | create_test | 0 | pass@1 | ↑ | 0.0671 | ± | 0.0196 | | |
| | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.4508 | ± | N/A | | |
| | | | none | 0 | inst_level_strict_acc | ↑ | 0.4269 | ± | N/A | | |
| | | | none | 0 | prompt_level_loose_acc | ↑ | 0.3142 | ± | 0.0200 | | |
| | | | none | 0 | prompt_level_strict_acc | ↑ | 0.2865 | ± | 0.0195 | | |
| | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | |
| | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | |
| --- | |
| ## Citation | |
| If you find our work helpful, please consider citing (citation will be updated after peer-reviewed publication): | |
| ```bibtex | |
| @misc{sinev-etal-2026-Zarya, | |
| author = {Sinev, Leonid and Koziev, Ilya and Leshchuk, Vladislav}, | |
| title = {Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference}, | |
| year = {2026}, | |
| archiveprefix = {arXiv}, | |
| eprint = {2609.19868}, | |
| primaryclass = {cs.CL}, | |
| url = {https://arxiv.org/abs/2609.19868}, | |
| } | |
| ``` | |