Instructions to use WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2") model = AutoModelForCausalLM.from_pretrained("WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2
- SGLang
How to use WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2 with Docker Model Runner:
docker model run hf.co/WooYoungSeok/qwen2.5-7b-instruct-student-likeness-error-generator-epoch2
Qwen2.5-7B-Instruct · student-likeness conditional error generator (GRPO, epoch 2.0)
Given a math problem and a misconception description, the model writes a student-style solution that
makes that error and ends with a (wrong) final answer. Trained with GRPO (TRL 1.14) from
Qwen/Qwen2.5-7B-Instruct. Run student_likeness_seed42_20260929_115310, snapshot epoch-2.0 (step 682 of 682,
the final model); it had the highest mean test reward of the five evaluated models.
Input format (use exactly this)
Chat template of Qwen2.5-Instruct with two messages:
- system: the text below with
{error_description}replaced by the misconception description - user: the problem text
You are a student who makes the following mathematical error.
Solve the problem in a way that naturally reflects this error.
Error description:
{error_description}
Show your working and give your final answer.
Write the answer itself, not an option letter or option number.
Write as that student would, believing your working is correct.
Do not mention the error, do not describe what you are doing wrong,
and do not state or compare with the correct answer.
Sampling used in training and evaluation: temperature 1.0, top_p 1.0, no top-k (0), repetition penalty 1.0,
max 1024 new tokens, stop on <|im_end|> / <|endoftext|>. generation_config.json holds exactly these settings,
so generate() and vLLM defaults sample like the RL run.
Training
- Data: Eedi misconception pairs (KEEP problems), split 80:20 by problem text; train 2,048 pairs / 965 questions.
- Reward per rollout = main + 0.5 × student-likeness + truncation
- main: +1 if the final answer is wrong (gpt-5-nano extraction and grading) and the reward verifier
(
WooYoungSeok/qwen2.5-math-7b-descriptive-verifier-v2-trval-halfA) labels both of 2 samples (T 0.6)alignedwith the misconception; 0 if wrong but not aligned; −0.75 if correct or no final answer - student-likeness: normalized pairwise win rate among the accepted rollouts of a group (gpt-5-nano judge, 2 real MathEDU student solutions as style references)
- truncation: −0.5 when cut off without EOS
- main: +1 if the final answer is wrong (gpt-5-nano extraction and grading) and the reward verifier
(
- GRPO: 8 rollouts per prompt, 48 rollouts (6 prompts) per optimizer step, 682 steps (2 epochs),
lr 1e-6 with 10% warmup and linear decay, beta 0.04, epsilon 0.2, loss
dapo, rewards scaled per group.
Test results (481 held-out pairs × 8 rollouts)
Scored like the training reward, but with the held-out verifier
WooYoungSeok/deepseek-r1-0528-qwen3-8b-descriptive-verifier-v2-trval-halfB. The snapshot was chosen on this
test set, so its scores are optimistic.
| metric | base Qwen2.5-7B-Instruct | this model |
|---|---|---|
| mean reward | 0.088 | 0.760 |
| success rate (wrong answer and both verifier samples aligned) | 41.8% | 75.0% |
| correct-answer rate | 57.4% | 23.6% |
| wrong answers accepted by the verifier | 98.1% | 98.2% |
| wrong answers equal to the target distractor | 24.7% | 21.2% |
Caveat: the test verifier accepts almost every wrong answer (base model included), and it separates misconceptions of the same question poorly. So the success rate mostly tracks how often the model answers wrongly, and it is weak evidence that the specific target misconception is followed.
- Downloads last month
- 5