WariHima/llm-jp-4-33b-thinking-Q4_K_M-GGUF

This model was converted to GGUF format from WariHima/llm-jp-4-33b-thinking using llama.cpp via the ggml.ai's GGUF-my-repo space. Refer to the original model card for more details on the model.

attention!!

llm-jp4のllm-jp-4-33b-thinkingをduplicateしていろいろしたリポジトリ(WariHima/llm-jp-4-33b-thinking) の物をgguf-my-repoで変換したモデルです。

chat templateをgpt-oss 11bから取ってくる(llm-jpの物を使う潜在的なバグ回避) tokenizerをllmjp4-tokenizerを使わないようtokenizer_config.jsonを変更 llm-jp4のmodels/ver4.0_alpha1.0/llm-jp-tokenizer_ver4.0_alpha1.0.modeleをtokenizer.modelにリネームし追加。

理論上はmainstreamのllama.cppでも動きますが、 しかし、ためしていないしうまく動かない可能性もあるので、 その際は、hiratagoh/llm-jp-4-32b-a3b-thinking-GGUFを参考になおしてください(丸投げします) またベンチマークもお任せします。

Downloads last month
256
GGUF
Model size
33B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WariHima/llm-jp-4-33b-thinking-Q4_K_M-GGUF

Quantized
(6)
this model