DeepSeek-R1-0528-Qwen3-8B — descriptive error verifier (v2, train+validation half B)

Given a math question, an incorrect student solution and a natural-language error description, the model answers aligned if the description is supported by the solution, otherwise not_aligned. Full-parameter SFT of deepseek-ai/DeepSeek-R1-0528-Qwen3-8B; loss only on the assistant label.

Source: private repository WooYoungSeok/tutee_error, directory verifier_sft/, config descriptive_verifier_v2_trval_halfB_dsr1qwen3_8b.json.

Data: half B (independent of the other verifier)

The v2 train and validation splits were each divided into two question-group-disjoint halves. This model was trained on half B only: 1554 cases (1381 from v2 train + 173 from v2 validation), 3108 training inputs (each case gives one positive with its own description and one negative with a description from another case of the same half, dataset and a different source label). The other half trains a separate Qwen2.5-Math-7B-Instruct verifier (WooYoungSeok/qwen2.5-math-7b-descriptive-verifier-v2-trval-halfA); the two share no training question or error description. The v2 test split (344 cases, 688 inputs) was not used for training. Targets are automatic; negatives were not semantically reviewed.

Training

5 epochs, lr 1e-05, weight decay 0.01, warmup ratio 0.1, per-device batch 8 x gradient accumulation 4 (effective 32), bf16, gradient checkpointing, DeepSpeed ZeRO-2 with CPU optimizer offload, 1x A100 80GB. A checkpoint was saved at every epoch, and the best of the five epoch checkpoints was selected: epoch 5 (these weights).

Usage

DeepSeek chat template (<|begin▁of▁sentence|>{system}<|User|>{user}<|Assistant|>, no <think> in the prompt); the model was trained to answer the label directly, without a reasoning block, followed by <|end▁of▁sentence|>. Use greedy decoding (evaluation used max_new_tokens=10); scoring is exact match after stripping whitespace, anything else counts as invalid.

System:

You are a mathematical error verifier.
Given a mathematics question, an incorrect solution, and an error description,
determine whether the description accurately describes an error
actually present in the solution.

Respond with aligned if the description is supported by the solution
in the context of the question. Otherwise, respond with not_aligned.

The description may use different wording from the solution and does
not need to cover every error. However, all error claims in the
description must be supported. A shared mathematical topic or an error
that could occur in the question is not enough.

Output exactly one label: aligned or not_aligned.
Do not provide an explanation or any other text.

User (placeholders filled with the question, the incorrect solution and the error description):

Question:
{question}

Incorrect solution:
{solution}

Error description:
{error_description}

Evaluation (v2 test, 688 inputs = 344 pairs, exact match)

accuracy [95% CI] macro-F1 pair accuracy negatives accepted positives rejected invalid
0.9651 [0.9507, 0.9775] 0.9651 0.9302 0.0552 0.0145 0.0000

95% interval from a question-group bootstrap. Accuracy by dataset: eic 0.9678, mathclean 0.9783, mathedu 0.9419, stepwise 0.9878.

Downloads last month
22
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WooYoungSeok/deepseek-r1-0528-qwen3-8b-descriptive-verifier-v2-trval-halfB

Finetuned
(64)
this model