|
Download README.md from atrost/climbmix-matformer-353m-1p2b-h100-m-vllm: direct link, hf CLI and curl.
- Browser
- Download file 1.2 kB
-
https://huggingface.co/atrost/climbmix-matformer-353m-1p2b-h100-m-vllm/resolve/main/README.md
- Command line
-
hf download hf://atrost/climbmix-matformer-353m-1p2b-h100-m-vllm/README.md
-
curl -L -o README.md https://huggingface.co/atrost/climbmix-matformer-353m-1p2b-h100-m-vllm/resolve/main/README.md
1.2 kB
metadata
license: cc-by-nc-4.0
language:
- en
datasets:
- nvidia/Nemotron-ClimbMix
tags:
- llama
- matformer
- matformer-m
- vllm
- causal-lm
- pretraining
- climbmix
atrost/climbmix-matformer-353m-1p2b-h100-m-vllm
Native LlamaForCausalLM conversion of atrost/climbmix-matformer-353m-1p2b-h100-m.
The source checkpoint was an extracted fixed-size MatFormer M submodel.
Because that submodel has fixed attention and MLP dimensions, its tensors can be
remapped losslessly into the standard Hugging Face Llama layout. This repo does
not require remote code and is intended to load through vLLM's native Llama path.
Source
- Source checkpoint:
atrost/climbmix-matformer-353m-1p2b-h100-m - Source revision:
cb222bb9353c30c372c7478533a5cc6f9feac1d7 - Local parity max logits diff before save:
0.000e+00
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "atrost/climbmix-matformer-353m-1p2b-h100-m-vllm"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForCausalLM.from_pretrained(repo_id, torch_dtype="auto")
from vllm import LLM
llm = LLM(model="atrost/climbmix-matformer-353m-1p2b-h100-m-vllm")