gemma-4-E4B-it-heretic-mlx-nvfp4

Converted from coder3101/gemma-4-E4B-it-heretic with mlx_vlm.convert.

  • Quantization: nvfp4
  • q_bits: 4
  • q_mode: nvfp4
  • Source type: Gemma 4 multimodal (Gemma4ForConditionalGeneration)
  • Converter env: /Users/vanch/.cache/lm-studio/.venvs/gemma4-mlx
  • Validation date: 2026-04-26

Validation Summary

  • Status: passed local text and vision smoke tests
  • Vision metadata present: processor_config.json and preprocessor_config.json
  • Peak memory across sampled tests: 7.669 GB

Validation Matrix

Case Output Prompt TPS Generation TPS Peak Memory
Text: one-sentence model description Gemma 4 is a family of open-weights large language models developed by Google DeepMind. 173.435 67.026 6.921 GB
Text: two-bullet structured answer * This model belongs to the Gemma family of open-weights models.
* As a member of this family, it can process and understand multimodal inputs, depending on the specific version being utilized.
238.298 65.270 6.933 GB
Vision: red image color identification The image is red. 576.505 73.123 7.627 GB
Vision: blue image color identification The image is blue. 570.144 73.028 7.669 GB

Local Artifacts

  • Full verification logs: /Users/vanch/models/gemma4-e4b-heretic-full-verify/nvfp4-*.log
Downloads last month
26
Safetensors
Model size
8B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vanch007/gemma-4-E4B-it-heretic-mlx-nvfp4

Quantized
(12)
this model