nvidia/Nemotron-ClimbMix
Viewer • Updated • 355M • 7.34k • 129
StairFormer trained on 1.2B tokens from nvidia/Nemotron-ClimbMix with nested auxiliary loss.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "atrost/climbmix-stairformer-353m-1p2b-h100"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
torch_dtype="auto",
)
trust_remote_code=True is required for the custom StairFormer/asymmetric checkpoints and harmless for the dense Llama baseline.
This clean repo was split out from atrost/climbmix-llama-288m-2p8b-h100 at commit a5b6981.