atrost's picture
Publish native Llama MatFormer M conversion from cb222bb9
f5ebd00 verified
|
Raw History Blame Contribute Delete
1.2 kB
metadata
license: cc-by-nc-4.0
language:
  - en
datasets:
  - nvidia/Nemotron-ClimbMix
tags:
  - llama
  - matformer
  - matformer-m
  - vllm
  - causal-lm
  - pretraining
  - climbmix

atrost/climbmix-matformer-353m-1p2b-h100-m-vllm

Native LlamaForCausalLM conversion of atrost/climbmix-matformer-353m-1p2b-h100-m.

The source checkpoint was an extracted fixed-size MatFormer M submodel. Because that submodel has fixed attention and MLP dimensions, its tensors can be remapped losslessly into the standard Hugging Face Llama layout. This repo does not require remote code and is intended to load through vLLM's native Llama path.

Source

  • Source checkpoint: atrost/climbmix-matformer-353m-1p2b-h100-m
  • Source revision: cb222bb9353c30c372c7478533a5cc6f9feac1d7
  • Local parity max logits diff before save: 0.000e+00

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "atrost/climbmix-matformer-353m-1p2b-h100-m-vllm"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, torch_dtype="auto")
from vllm import LLM

llm = LLM(model="atrost/climbmix-matformer-353m-1p2b-h100-m-vllm")