--- license: apache-2.0 base_model: Qwen/Qwen3-4B-Base tags: - reinforcement-learning - grpo - code-generation - verl - easyr1 datasets: - agentica-org/DeepCoder-Preview-Dataset language: - en pipeline_tag: text-generation --- # Qwen3-4B-Base — TACO code RL (baseline) RL-finetuned from **Qwen/Qwen3-4B-Base** with **GRPO** on the `taco` subset of [agentica-org/DeepCoder-Preview-Dataset](https://huggingface.co/datasets/agentica-org/DeepCoder-Preview-Dataset), using an **execution-based reward** (generated programs are run against the problem's stdin/stdout test cases). **Method:** standard GRPO (no resampling) ## Results (pass@1, %) | Benchmark | This model | Qwen3-4B-Base (before RL) | |---|---|---| | HumanEval | **84.15** | 76.83 | | MBPP | **60.60** | 49.00 | | Average | **72.38** | 62.91 | All four method variants trained under identical settings: | Method | HumanEval | MBPP | Average | |---|---|---|---| | baseline (GRPO) | 84.15 | 60.60 | 72.38 | | naive resample | 84.76 | 57.40 | 71.08 | | **lorem** | **85.98** | 66.00 | **75.99** | | lorem + shaping | 81.71 | 66.20 | 73.96 | Evaluated with [evalscope](https://github.com/modelscope/evalscope), pass@1, single sample, temperature=0, served via vLLM. Note: training prompts asked the model to reason inside `` tags, while the evaluation used the benchmarks' default (non-think) prompts. ## Training setup - **Framework:** [verl](https://github.com/volcengine/verl) / EasyR1 fork - **Steps:** 174 (3 epochs over 7432 problems) - **Rollout:** n=8 per prompt, batch size 128 - **Max length:** 2048 prompt / 8192 response - **KL:** disabled - **Hardware:** 8×A100 (single node) ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "shrango/qwen3-4b-base-taco-grpo-baseline" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto") ```