Bug Report

#1
by rinoa - opened

llama.cpp (version: 6708 (df1b612e2)) fails to recognize this gguf model:

llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'lfm2moe'
llama_model_load_from_file_impl: failed to load model
common_init_from_params: failed to load model 'LFM2-8B-A1B-Q5_K_M.gguf', try reducing --n-gpu-layers if you're running out of VRAM
srv    load_model: failed to load model, 'LFM2-8B-A1B-Q5_K_M.gguf'
srv    operator(): operator(): cleaning up before exit...
main: exiting due to model loading error
Liquid AI org

Upstream PR is still in review https://github.com/ggml-org/llama.cpp/pull/16464. Once it's merged, please download the updated binaries and launch them. Or compile binaries from the PR.

tdakhran changed discussion status to closed
Liquid AI org

Latest llama.cpp release supports lfm2-moe https://github.com/ggml-org/llama.cpp/releases/tag/b6709!

I have just pull latest release (b6710) from llama.cpp and i get this error when loading the model:

error loading model: missing tensor 'blk.2.exp_probs_b.bias'

++

Victor

Liquid AI org

@vico44 GGUFs were reuploaded yesterday. Please redownload and verify that the SHA256 sum matches the one shown here: https://huggingface.co/LiquidAI/LFM2-8B-A1B-GGUF/blob/main/LFM2-8B-A1B-Q6_K.gguf. Make sure you select the right quantization.

Sign up or log in to comment