How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker
docker model run hf.co/WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2
Quick Links

Qwen2.5-7B-Instruct · student-likeness conditional error generator (GRPO, epoch 2.0)

Given a math problem and a misconception description, the model writes a student-style solution that makes that error and ends with a (wrong) final answer. Trained with GRPO (TRL 1.14) from Qwen/Qwen2.5-7B-Instruct. Run student_likeness_seed42_20260929_115310, snapshot epoch-2.0 (step 682 of 682, the final model); it had the highest mean test reward of the five evaluated models.

Input format (use exactly this)

Chat template of Qwen2.5-Instruct with two messages:

  • system: the text below with {error_description} replaced by the misconception description
  • user: the problem text
You are a student who makes the following mathematical error.
Solve the problem in a way that naturally reflects this error.

Error description:
{error_description}

Show your working and give your final answer.
Write the answer itself, not an option letter or option number.
Write as that student would, believing your working is correct.
Do not mention the error, do not describe what you are doing wrong,
and do not state or compare with the correct answer.

Sampling used in training and evaluation: temperature 1.0, top_p 1.0, no top-k (0), repetition penalty 1.0, max 1024 new tokens, stop on <|im_end|> / <|endoftext|>. generation_config.json holds exactly these settings, so generate() and vLLM defaults sample like the RL run.

Training

  • Data: Eedi misconception pairs (KEEP problems), split 80:20 by problem text; train 2,048 pairs / 965 questions.
  • Reward per rollout = main + 0.5 × student-likeness + truncation
    • main: +1 if the final answer is wrong (gpt-5-nano extraction and grading) and the reward verifier (WooYoungSeok/qwen2.5-math-7b-descriptive-verifier-v2-trval-halfA) labels both of 2 samples (T 0.6) aligned with the misconception; 0 if wrong but not aligned; −0.75 if correct or no final answer
    • student-likeness: normalized pairwise win rate among the accepted rollouts of a group (gpt-5-nano judge, 2 real MathEDU student solutions as style references)
    • truncation: −0.5 when cut off without EOS
  • GRPO: 8 rollouts per prompt, 48 rollouts (6 prompts) per optimizer step, 682 steps (2 epochs), lr 1e-6 with 10% warmup and linear decay, beta 0.04, epsilon 0.2, loss dapo, rewards scaled per group.

Test results (481 held-out pairs × 8 rollouts)

Scored like the training reward, but with the held-out verifier WooYoungSeok/deepseek-r1-0528-qwen3-8b-descriptive-verifier-v2-trval-halfB. The snapshot was chosen on this test set, so its scores are optimistic.

metric base Qwen2.5-7B-Instruct this model
mean reward 0.088 0.760
success rate (wrong answer and both verifier samples aligned) 41.8% 75.0%
correct-answer rate 57.4% 23.6%
wrong answers accepted by the verifier 98.1% 98.2%
wrong answers equal to the target distractor 24.7% 21.2%

Caveat: the test verifier accepts almost every wrong answer (base model included), and it separates misconceptions of the same question poorly. So the success rate mostly tracks how often the model answers wrongly, and it is weak evidence that the specific target misconception is followed.

Downloads last month
7
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2

Base model

Qwen/Qwen2.5-7B
Finetuned
(3142)
this model