How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "modrill/qwen3-4b-think-s1-ep23-full-sft"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "modrill/qwen3-4b-think-s1-ep23-full-sft",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker
docker model run hf.co/modrill/qwen3-4b-think-s1-ep23-full-sft
Quick Links

Qwen3-4B Think S1 Ep23 (Full SFT)

Full-parameter supervised fine-tuning (SFT) of Qwen/Qwen3-4B-Base on the ocr_think_50k dataset with the Qwen3 chat template (think-style reasoning).

This checkpoint is stage 1 episode 23: training continued from an internal think_s1 run at checkpoint-302 (same base architecture), then fine-tuned for two additional epochs on ocr_think_50k.

Model description

  • Method: full SFT (all weights trainable), DeepSpeed ZeRO-3, 4 GPUs
  • Dataset: ocr_think_50k
  • Template: qwen3
  • Not LoRA / not QLoRA: entire 4B model was updated

Training details

Field Value
Epochs 2
Seed 42
cutoff_len 24576
packing true
neat_packing false
per_device_train_batch_size 1
gradient_accumulation_steps 16
effective_batch_size 64
learning_rate 5e-5
train_loss 0.5416
train_steps 604
finished_at 2026-06-10 05:23 CST

Optimizer: AdamW (fused), cosine schedule, warmup ratio 0.1. Framework: Transformers 5.6.0, PyTorch 2.8.0+cu128.

Related models

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "modrill/qwen3-4b-think-s1-ep23-full-sft"
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)

License

Released under Apache 2.0 (see LICENSE in the upstream Qwen model card if not bundled here).

Downloads last month
9
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for modrill/qwen3-4b-think-s1-ep23-full-sft

Finetuned
(423)
this model