gemma-4-12B-TNG-V8

SkullOfStars

This is a Nightmedia release, trained with locally generated content.

This training includes TNG-infused coding lessons.

The teacher model used was the qx86-hi quant of the nightmedia/Qwen3.6-35B-A3B-Holo3-Qwopus-bf16

The lessons are exclusively in backend engineering, LLM design and training, operational Holodeck in Haskell, Rust, Golang and Python.

The set up was in the DS9 Holodeck, with in-person commentary by Spock, Data, Quark, Worf, Odo, Julian, Garak, Q, and many others.

Additionally, some of the Polaris Alpha questions were used to distill from the teacher model.

         arc   arc/e boolq hswag obkqa piqa  wino
bf16     0.553,0.771,0.849
mxfp8    0.551,0.779,0.865
q8-hi    0.552,0.770,0.848
qx86-hi  0.539,0.771,0.835
mxfp4    0.505,0.750,0.853

Quant    Perplexity      Peak Memory   Tokens/sec
bf16     67.645 ยฑ 1.080  30.95 GB      549
mxfp8    74.437 ยฑ 1.177  19.42 GB      373
q8-hi    65.357 ยฑ 1.031  20.54 GB      460
qx86-hi  64.724 ยฑ 1.021  19.00 GB      394
mxfp4   117.640 ยฑ 2.029  13.47 GB      442

Training TNG base

gemma-4-12B-TNG-V6-q8-hi-mlx

         arc   arc/e boolq hswag obkqa piqa  wino
q8-hi    0.524,0.713,0.775

Quant    Perplexity      Peak Memory   Tokens/sec
q8-hi    75.792 ยฑ 1.251  20.54 GB      446

Training 1/4 curriculum

gemma-4-12B-TNG-V1-q8-hi-mlx

         arc   arc/e boolq hswag obkqa piqa  wino
mxfp8    0.439,0.617,0.695
q8-hi    0.475,0.615,0.724
qx86-hi  0.468,0.614,0.753

Quant    Perplexity      Peak Memory   Tokens/sec
mxfp8    85.212 ยฑ 1.319  19.42 GB      360
q8-hi    73.981 ยฑ 1.145  20.54 GB      373
qx86-hi  71.764 ยฑ 1.103  19.00 GB      416

Baseline model

gemma-4-12B-it

         arc   arc/e boolq hswag obkqa piqa  wino
mxfp8    0.385,0.527,0.766,0.509,0.386,0.664,0.579
q8-hi    0.394,0.522,0.774,0.520,0.370,0.674,0.583
qx86-hi  0.391,0.526,0.777,0.521,0.366,0.677,0.590
mxfp4    0.371,0.517,0.660,0.508,0.368,0.677,0.581

Quant    Perplexity      Peak Memory   Tokens/sec
mxfp8   175.766 ยฑ 3.092  19.42 GB      463
mxfp4   283.801 ยฑ 5.361  13.47 GB      484

The V8 iteration introduces a random pick of Claude code traces from angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k-raw mixed in with a random pick of lessons from the first stage.

-G

Downloads last month
3,519
Safetensors
Model size
12B params
Tensor type
BF16
ยท
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for nightmedia/gemma-4-12B-TNG-V8

Finetuned
(188)
this model
Quantizations
2 models

Space using nightmedia/gemma-4-12B-TNG-V8 1