Llama-3.2 3b Instruct EAGLE3 ShareGPT SW512

This repository contains a SpecForge EAGLE3 draft-model checkpoint for use with meta-llama/Llama-3.2-3B-Instruct. It is a draft model for speculative decoding, not a standalone target language model.

Checkpoint

Field Value
Source run llama3.2-3b-inst-eagle3-sharegpt-sw512
Checkpoint epoch_9_step_231810
Epoch 9
Global step 231810
Files config.json, model.safetensors, training_state.pt (when present)

Training Parameters

Parameter Value
Base model meta-llama/Llama-3.2-3B-Instruct
Method SpecForge EAGLE3 online training
Training data sharegpt_train_clean.jsonl
Learning rate 0.0001
Batch size 1
Target batch size 1
Epochs configured 10
Max length 2048
Warmup ratio 0.015
Max grad norm 0.5
TTT length 7
Draft layers Not recorded
Draft sliding window 512
Save interval 5000
Eval interval 5000
Seed 0
TP / DP size 1 / 4
Attention backend sdpa
Target model backend sglang
SGLang attention backend flashinfer
Dataset build workers 64

Draft Model Configuration

Field Value
Architecture LlamaForCausalLMEagle3
dtype bfloat16
Hidden size 3072
Intermediate size 8192
Draft layers 1
Attention heads 24
KV heads 8
Draft vocab size 32000
Vocab size 128256
Max position embeddings 131072
Sliding window 512
Max window layers Not recorded
Future hidden Not recorded

Notes

  • This is the highest-step local checkpoint available when this repository was published.
  • The checkpoint is intended to be loaded by SpecForge/EAGLE3-compatible code.
  • training_state.pt is included when available for provenance and training-state inspection.
  • No benchmark claim is made in this card.
Downloads last month
346
Safetensors
Model size
0.2B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for huluhuluu/llama3p2-3b-inst-eagle3-sharegpt-sw512-epoch-9-step-231810

Finetuned
(1999)
this model

Collection including huluhuluu/llama3p2-3b-inst-eagle3-sharegpt-sw512-epoch-9-step-231810