--- license: mit language: - ru - en pipeline_tag: text-generation tags: - dllm - diffusion - diffusion-language-modeling - instruct library_name: transformers --- # Zarya-1.7B Zarya is a family of hybrid language models that combine a classic auto-regressive (AR) objective with a masked-diffusion (MDM) objective in one model. The architecture can be built on top of any autoregressive model but in this repository it uses the `Qwen3` backbone. Naming explanation: Zarya (pronounced as [zɐˈrʲa] ([IPA notation](https://en.wiktionary.org/wiki/Appendix:Russian_pronunciation)), literally "Dawn" in English) is a figure from Slavic folklore — a female personification of dawn who may be considered a goddess. In various traditions, she can manifest as a single being or as two or three sisters simultaneously. This is a research prototype. ## Model Details ### Model Description Zarya is a research prototype of a family of hybrid language models that jointly learn a classic auto-regressive (AR) objective and a masked-diffusion (MDM) objective within a single model. Two generation modes are supported, both reachable through a single `model.generate(...)` call: masked-diffusion (MDM) sampling and slotted-level speculative parallel decoding. - **Model type:** Hybrid auto-regressive (AR) + masked-diffusion language model (DLLM); backbone `Qwen3`, wrapper `Zarya` - **Language(s) (NLP):** Russian and English - **License:** MIT - **Preprint:** https://arxiv.org/abs/2609.19868 - **Repository with training code:** https://github.com/ai-forever/zarya ### Zarya-1.7B details Zarya-1.7B has the following features: | Variant | hidden_size | num_hidden_layers | num_attention_heads | intermediate_size | |------------|-------------|-------------------|---------------------|-------------------| | Zarya-1.7B | 2048 | 28 | 16 | 6144 | Context Length: 2048 ## Uses Zarya is intended for text generation. It supports conversational fine-tuning (SFT) and classic auto-regressive pretraining. ### Direct Use Direct use is text generation (continuation of a prompt) through the `model.generate(...)` interface, including chat-style prompts formatted with the provided chat template. Two inference modes are available through the same `generate()` call. Both modes fully use the KV cache with causal attention masks. - **MDM sampling** (`slotted_generation=false`): iterative masked-diffusion denoising with the first-hitting sampler. - **Slotted speculative decoding** (`slotted_generation=true`): parallel slot generation with inter-slot diffusion-based selection and intra-slot autoregressive generation for a decoding speedup. ### Out-of-Scope Use The model is a research prototype. It should not be used for production decisions, safety-critical applications, or any use case where accuracy and reliability are essential without additional evaluation and safeguards. Inference performance and stability also depend on the chosen decoding hyperparameters (like `slotted_generation`, `slot_size`, `serial_num_blocks`, `slot_threshold`, `token_threshold`, and others). ## Bias, Risks, and Limitations This is a research prototype. The code relies on Hugging Face Transformers APIs; when upgrading versions, compatibility must be checked (tested on Transformers 5.12.1 and PyTorch 2.9.0). Inference performance and stability depend on the choice of config parameters. ## How to Get Started with the Model Use the code below to get started with the model. Loading the model and tokenizer requires `trust_remote_code=True`. ```python import torch from transformers import AutoModel, AutoTokenizer model_name = "ai-forever/Zarya-1.7B" model = AutoModel.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.bfloat16).cuda() tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True) prompt = "<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n" input_ids = tokenizer(prompt, return_tensors="pt").input_ids.cuda() # Both modes go through model.generate(...). # With generation_config.slotted_generation=true -> slotted speculative decoding: out = model.generate( input_ids, max_new_tokens=256, do_sample=True, temperature=0.7, slot_size=16, serial_num_blocks=4, slot_threshold=0.9, token_threshold=0.3, ) # Setting generation_config.slotted_generation=false -> MDM sampling instead: # out = model.generate(input_ids, max_new_tokens=256) print(tokenizer.decode(out[0])) ``` ## Evaluation LM-eval benchmarking with the `lm-eval` package is supported. Example run: ```bash lm_eval run \ --tasks=gsm8k,ifeval,mbpp,mbpp_instruct,mbpp_plus,mbpp_plus_instruct,hellaswag \ --model=hf --confirm_run_unsafe_code \ --log_samples \ --apply_chat_template \ --output_path=./reports/lm-eval_results \ --model_args=pretrained=ai-forever/Zarya-1.7B,backend=causal,dtype=bfloat16,attn_implementation=sdpa,trust_remote_code=True \ --gen_kwargs slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4 ``` ### Zarya-1.7B results Measurements below were collected with varying inference parameters and on different GPUs; performance is sensitive to both, so results may differ across configurations and hardware setups. #### A100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4` Hardware info: gpu_driver_cuda_version 13.2; gpu_driver_version 595.71.05; Docker info: Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; `lm-eval == 0.4.12` | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | |-----------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.4670 | ± | 0.0137 | | | | strict-match | 5 | exact_match | ↑ | 0.4685 | ± | 0.0137 | | hellaswag | 1 | none | 0 | acc | ↑ | 0.4224 | ± | 0.0049 | | | | none | 0 | acc_norm | ↑ | 0.5472 | ± | 0.0050 | | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.5815 | ± | N/A | | | | none | 0 | inst_level_strict_acc | ↑ | 0.5552 | ± | N/A | | | | none | 0 | prompt_level_loose_acc | ↑ | 0.4492 | ± | 0.0214 | | | | none | 0 | prompt_level_strict_acc | ↑ | 0.4288 | ± | 0.0213 | | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.3540 | ± | 0.0214 | | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.4815 | ± | 0.0257 | #### H100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4` Hardware info: gpu_driver_cuda_version 13.0; gpu_driver_version 580.105.08; Docker info: Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; `lm-eval == 0.4.12` | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | |--------------------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | arc_challenge | 1 | none | 0 | acc | ↑ | 0.4437 | ± | 0.0145 | | | | none | 0 | acc_norm | ↑ | 0.4573 | ± | 0.0146 | | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.0553 | ± | 0.0063 | | | | strict-match | 5 | exact_match | ↑ | 0.0455 | ± | 0.0057 | | hellaswag | 1 | none | 0 | acc | ↑ | 0.4227 | ± | 0.0049 | | | | none | 0 | acc_norm | ↑ | 0.5474 | ± | 0.0050 | | hendrycks_math500 | 1 | none | 0 | exact_match | ↑ | 0.0040 | ± | 0.0028 | | humaneval | 1 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | humaneval_instruct | 4 | create_test | 0 | pass@1 | ↑ | 0.0549 | ± | 0.0178 | | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.4317 | ± | N/A | | | | none | 0 | inst_level_strict_acc | ↑ | 0.4053 | ± | N/A | | | | none | 0 | prompt_level_loose_acc | ↑ | 0.3013 | ± | 0.0197 | | | | none | 0 | prompt_level_strict_acc | ↑ | 0.2791 | ± | 0.0193 | | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | #### H100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=false,T=0,temperature=0.5,top_p=0.8,do_sample=true,noise_schedule=linear` Hardware info: gpu_driver_cuda_version 13.0; gpu_driver_version 580.105.08; Docker info: Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; `lm-eval == 0.4.12` | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | |--------------------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | arc_challenge | 1 | none | 0 | acc | ↑ | 0.4437 | ± | 0.0145 | | | | none | 0 | acc_norm | ↑ | 0.4573 | ± | 0.0146 | | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.0091 | ± | 0.0026 | | | | strict-match | 5 | exact_match | ↑ | 0.0000 | ± | 0 | | hellaswag | 1 | none | 0 | acc | ↑ | 0.4227 | ± | 0.0049 | | | | none | 0 | acc_norm | ↑ | 0.5474 | ± | 0.0050 | | hendrycks_math500 | 1 | none | 0 | exact_match | ↑ | 0.0000 | ± | 0 | | humaneval | 1 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | humaneval_instruct | 4 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.2002 | ± | N/A | | | | none | 0 | inst_level_strict_acc | ↑ | 0.1655 | ± | N/A | | | | none | 0 | prompt_level_loose_acc | ↑ | 0.1072 | ± | 0.0133 | | | | none | 0 | prompt_level_strict_acc | ↑ | 0.0776 | ± | 0.0115 | | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | #### H100, `dtype=bfloat16`, `apply_chat_template`, `slotted_generation=true,slot_size=16,serial_num_blocks=4,slot_threshold=0.9,token_threshold=0.4,max_gen_toks=2048` Hardware info: gpu_driver_cuda_version 13.0; gpu_driver_version 580.105.08; Docker info: Torch: 2.9.0+cu128; Transformers: 5.12.1; CUDNN in torch: 91002; `lm-eval == 0.4.12` | Tasks | Version | Filter | n-shot | Metric | | Value | | Stderr | |--------------------|--------:|------------------|-------:|-------------------------|---|-------:|---|--------| | arc_challenge | 1 | none | 0 | acc | ↑ | 0.4437 | ± | 0.0145 | | | | none | 0 | acc_norm | ↑ | 0.4573 | ± | 0.0146 | | gsm8k | 3 | flexible-extract | 5 | exact_match | ↑ | 0.0167 | ± | 0.0035 | | | | strict-match | 5 | exact_match | ↑ | 0.0099 | ± | 0.0027 | | hellaswag | 1 | none | 0 | acc | ↑ | 0.4227 | ± | 0.0049 | | | | none | 0 | acc_norm | ↑ | 0.5474 | ± | 0.0050 | | hendrycks_math500 | 1 | none | 0 | exact_match | ↑ | 0.0000 | ± | 0 | | humaneval | 1 | create_test | 0 | pass@1 | ↑ | 0.0000 | ± | 0 | | humaneval_instruct | 4 | create_test | 0 | pass@1 | ↑ | 0.0671 | ± | 0.0196 | | ifeval | 4 | none | 0 | inst_level_loose_acc | ↑ | 0.4508 | ± | N/A | | | | none | 0 | inst_level_strict_acc | ↑ | 0.4269 | ± | N/A | | | | none | 0 | prompt_level_loose_acc | ↑ | 0.3142 | ± | 0.0200 | | | | none | 0 | prompt_level_strict_acc | ↑ | 0.2865 | ± | 0.0195 | | mbpp | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | | mbpp_plus | 1 | none | 3 | pass_at_1 | ↑ | 0.0000 | ± | 0 | --- ## Citation If you find our work helpful, please consider citing (citation will be updated after peer-reviewed publication): ```bibtex @misc{sinev-etal-2026-Zarya, author = {Sinev, Leonid and Koziev, Ilya and Leshchuk, Vladislav}, title = {Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference}, year = {2026}, archiveprefix = {arXiv}, eprint = {2609.19868}, primaryclass = {cs.CL}, url = {https://arxiv.org/abs/2609.19868}, } ```