Considerably better than the other abliterated model flying around

#1
by wacomctl672 - opened

Had quantized the model to gguf and used the audio vision mmproj (from ggml repository, I don't know whether this makes a difference but it worked) and it works well at q4km (maybe not as relevant since mlx focused). Thank you for the abliteration! edit: had to make an edit to llama.cpp\conversion\nemotron.py and in it def modify_tensors line 431, which i forgot to mention. i added this to the beginning of the function:

if "switch_mlp.fc1.weight" in name:
            yield f"blk.{bid}.ffn_up_exps.weight", data_torch
            return
elif "switch_mlp.fc2.weight" in name:
            yield f"blk.{bid}.ffn_down_exps.weight", data_torch
            return

thanks finn, glad it's holding up for you. good to know the ggml audio vision mmproj works with the gguf, that's useful for anyone on llama.cpp instead of mlx. i'll add a note about it in the model card.

uploaded it to save somebody the time if it ever gets needed. will take down if you dont like it. https://huggingface.co/wacomctl672/Nemotron-3-Nano-Omni-30B-Abliterated-MM-GGUF

thanks for doing that, definitely keep it up. i only did the mlx side so a gguf fills a real gap for the llama.cpp folks, and good to know the ggml mmproj works with it. appreciate you sharing it back instead of sitting on it.

thanks
matt

Sign up or log in to comment