Instructions to use nightmedia/gemma-4-12B-TNG-V8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nightmedia/gemma-4-12B-TNG-V8 with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("nightmedia/gemma-4-12B-TNG-V8") model = AutoModelForMultimodalLM.from_pretrained("nightmedia/gemma-4-12B-TNG-V8", device_map="auto") - MLX
How to use nightmedia/gemma-4-12B-TNG-V8 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir gemma-4-12B-TNG-V8 nightmedia/gemma-4-12B-TNG-V8
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
gemma-4-12B-TNG-V8
This is a Nightmedia release, trained with locally generated content.
This training includes TNG-infused coding lessons.
The teacher model used was the qx86-hi quant of the nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-bf16
The lessons are exclusively in backend engineering, LLM design and training, operational Holodeck in Haskell, Rust, Golang and Python.
The set up was in the DS9 Holodeck, with in-person commentary by Spock, Data, Quark, Worf, Odo, Julian, Garak, Q, and many others.
Additionally, some of the Polaris Alpha questions were used to distill from the teacher model.
arc arc/e boolq hswag obkqa piqa wino
bf16 0.553,0.771,0.849
mxfp8 0.551,0.779,0.865
q8-hi 0.552,0.770,0.848
qx86-hi 0.539,0.771,0.835
mxfp4 0.505,0.750,0.853
Quant Perplexity Peak Memory Tokens/sec
bf16 67.645 ยฑ 1.080 30.95 GB 549
mxfp8 74.437 ยฑ 1.177 19.42 GB 373
q8-hi 65.357 ยฑ 1.031 20.54 GB 460
qx86-hi 64.724 ยฑ 1.021 19.00 GB 394
mxfp4 117.640 ยฑ 2.029 13.47 GB 442
Training TNG base
gemma-4-12B-TNG-V6-q8-hi-mlx
arc arc/e boolq hswag obkqa piqa wino
q8-hi 0.524,0.713,0.775
Quant Perplexity Peak Memory Tokens/sec
q8-hi 75.792 ยฑ 1.251 20.54 GB 446
Training 1/4 curriculum
gemma-4-12B-TNG-V1-q8-hi-mlx
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.439,0.617,0.695
q8-hi 0.475,0.615,0.724
qx86-hi 0.468,0.614,0.753
Quant Perplexity Peak Memory Tokens/sec
mxfp8 85.212 ยฑ 1.319 19.42 GB 360
q8-hi 73.981 ยฑ 1.145 20.54 GB 373
qx86-hi 71.764 ยฑ 1.103 19.00 GB 416
Baseline model
gemma-4-12B-it
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.385,0.527,0.766,0.509,0.386,0.664,0.579
q8-hi 0.394,0.522,0.774,0.520,0.370,0.674,0.583
qx86-hi 0.391,0.526,0.777,0.521,0.366,0.677,0.590
mxfp4 0.371,0.517,0.660,0.508,0.368,0.677,0.581
Quant Perplexity Peak Memory Tokens/sec
mxfp8 175.766 ยฑ 3.092 19.42 GB 463
mxfp4 283.801 ยฑ 5.361 13.47 GB 484
The V8 iteration introduces a random pick of Claude code traces from angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k-raw mixed in with a random pick of lessons from the first stage.
-G
- Downloads last month
- 3,519
Quantized
