--- base_model: nex-agi/Nex-N2-mini library_name: transformers pipeline_tag: text-generation license: apache-2.0 tags: - qwen3.6 - nex - gptq - gptq-pro - foem - moe - marlin - vllm - int4 - quantized - long-context - tool-use - function-calling - terminal-bench - non-mtp - multimodal-tensors-preserved ---  # Nex-N2-mini GPTQ-Pro
These models are built and maintained on rented GPU compute. If you want to show some appreciation, a follow on X or a coffee helps keep the releases coming.
This is a GPTQ-Pro 4-bit quantization of [`nex-agi/Nex-N2-mini`](https://huggingface.co/nex-agi/Nex-N2-mini). It is a deployment artifact, not a new fine-tune. The goal is to make the Nex-N2-mini MoE checkpoint easier to test in GPTQ-compatible local serving stacks while keeping the model card honest about the validation status. The source checkpoint includes vision/visual tensors. This artifact preserves those tensors, but the validated publication story here is text and coding-agent serving. Vision behavior has not yet been validated for the quantized artifact. ## Source And Credits Source model: - [`nex-agi/Nex-N2-mini`](https://huggingface.co/nex-agi/Nex-N2-mini) Quantization tooling and reference recipe: - [`groxaxo/GPTQ-Pro`](https://github.com/groxaxo/GPTQ-Pro) - [`groxaxo/Qwen3.6-27B-GPTQ-Pro-4bit`](https://huggingface.co/groxaxo/Qwen3.6-27B-GPTQ-Pro-4bit) ## Artifact Summary | Field | Value | |---|---:| | Source model | `nex-agi/Nex-N2-mini` | | Architecture | `Qwen3_5MoeForConditionalGeneration` | | Model type | `qwen3_5_moe` | | Tensor files | `5` | | Safetensors size | `19.23 GiB` | | Indexed tensors | `124576` | | Quantized `qweight` tensors | `30970` | | `mtp.*` tensors in index | `false` | | vision/visual tensors in index | `true` | | Index metadata size matches shards | `true` | The source index/logs showed no `mtp.*` tensors. This artifact therefore normalizes `text_config.mtp_num_hidden_layers` to `0` and records the change under `artifact_notes.mtp`. ## Quantization Recipe | Setting | Value | |---|---:| | Method | GPTQ-Pro / GPTQModel | | Quantizer | `gptqmodel:6.1.0-dev` | | Bits | `4` | | Group size | `128` | | Symmetric quantization | `true` | | Desc act | `false` | | True sequential | `true` | | Calibration dataset | WikiText | | Calibration samples | `256` | | Calibration sequence length | `2048` | | MSE | `2.0` | | Damp percent | `0.05` | | Damp auto increment | `0.01` | | FOEM alpha | `0.25` | | FOEM beta | `0.2` | | FOEM device | `cuda:0` | | MoE routing | `ExpertsRoutingBypass` | | MoE bypass batch size | `320` | | Dense VRAM strategy | `exclusive` | | MoE VRAM strategy | `balanced` | | Pack implementation | `cpu` | Fallback smoothing was enabled for difficult groups with threshold `0.5%`. ## Intended Serving Shape This checkpoint is intended for advanced users testing text-only GPTQ serving for Qwen3.6-style MoE models. A starting vLLM shape for text-only testing: ```bash vllm serve XReyRobert/Nex-N2-mini-GPTQ-Pro \ --served-model-name nex-n2-mini-gptq-pro \ --language-model-only \ --dtype float16 \ --quantization gptq_marlin \ --tensor-parallel-size 1 \ --max-model-len 262144 \ --max-num-seqs 1 \ --kv-cache-dtype fp8_e5m2 \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --enable-prefix-caching \ --gpu-memory-utilization 0.95 \ --trust-remote-code ``` Serving context for the published Smoke24/vLLM measurements: The Smoke24/vLLM numbers were collected on an internal llm-residency vLLM deployment. The custom image recipe is not published yet, so this card does not present that image as a public reproduction target. The stable serving knobs captured from the run are listed for context. | Field | Value | |---|---| | Nomad job profile | `vllm-nex-n2-mini-262k` | | Served model name | `nex-n2-mini-gptq-pro-ctx262k` | | Critical flags | `--dtype float16`, `--quantization gptq_marlin`, `--kv-cache-dtype fp8_e5m2`, `--reasoning-parser qwen3`, `--tool-call-parser qwen3_coder`, `--max-model-len 262144`, `--max-num-batched-tokens 2096` | | Benchmark context | Smoke24 quality rows used `max_model_len=131072` for apples-to-apples comparison; the image above reflects the validated 262k residency serving profile. | Treat this as a starting point. Loader compatibility depends on vLLM, Transformers, GPTQModel, GPTQ-Marlin, and Qwen3.6 MoE support. The RTX 3090 image above reflects separate 262k-context serving validation. ## Public vLLM Reproducibility This artifact has a public reproducibility path on the unmodified upstream vLLM OpenAI image: - image: `docker.io/vllm/vllm-openai:nightly-7a1eb8ac2ec4ea69338c51dc7afd4b15010abfa8` - vLLM version observed in validation: `0.20.1rc1.dev16+g7a1eb8ac2` - GPU class: single RTX 3090 24 GB / Ampere - `--enforce-eager` was **not** used - no local sleep/wake patch or `localhost/*sleepwake*` image is required for the validation below Validated serving shapes: - `--max-model-len 131072` validated with `--gpu-memory-utilization 0.94` - `--max-model-len 262144` validated with `--gpu-memory-utilization 0.96` - `--language-model-only`, `--dtype float16`, `--quantization gptq_marlin` - `--kv-cache-dtype fp8_e5m2`, `--enable-prefix-caching`, `--max-num-seqs 1` - `--max-num-batched-tokens 2096`, `--max-cudagraph-capture-size 32` - `--reasoning-parser qwen3`, `--tool-call-parser qwen3_coder` The 262k profile is tight on 24 GB GPUs; the lower `0.94` memory target used for 131k was short on KV cache, while `0.96` passed. ## Validation And Benchmarks Completed artifact checks: - Local shard index inspection completed before upload. - Remote file list verified after upload. - Remote `model.safetensors.index.json` verified after upload. - Index metadata total size matches the local safetensor shards. - The remote artifact contains the expected five safetensor shards. Terminal-Bench 2.0 Smoke24 result and associated vLLM serving measurements. This Smoke24 run used `max_model_len=131072` for apples-to-apples comparison with the other local models in this publication batch: | Run | Score | Success rate | Wall-time | Output tokens | Observed decode | LLM API time | |---|---:|---:|---:|---:|---:|---:| | `nex-n2-mini-gptq-pro` | `14/24` | `58.3%` | `314.6m` | `1670.6k` | `140.8 tok/s` | `197.4m` | Smoke24 is a fixed 24-task Terminal-Bench 2.0 comparison corpus, not a full Terminal-Bench leaderboard run. In this harness, Nex-N2-mini GPTQ-Pro tied the Qwen3.6 27B GPTQ reference on solved tasks but used more wall time and far more output tokens. That makes it a useful candidate for further serving and generation-control tuning, not an efficiency leader in this specific test. Task list and harness shape: - [`benchmarks/terminal-bench-2.0/smoke24_task_list_20260616.md`](benchmarks/terminal-bench-2.0/smoke24_task_list_20260616.md) ## MTP And Vision Status - `mtp.*` tensors are not present in this artifact. - `text_config.mtp_num_hidden_layers` was normalized to `0`. - Do not enable MTP speculative decoding for this artifact. - Vision/visual tensors are present, but multimodal serving has not been validated for this quantized artifact. ## Limitations - Experimental quantization. - Terminal-Bench Smoke24 is a small local comparison corpus, not a full benchmark submission. - Nex-N2-mini was verbose and reasoning-heavy in the Smoke24 harness; generation controls may need further tuning. - MTP speculative decoding is not supported by this artifact. - Vision tensors are preserved, but vision behavior has not been validated. - Loader behavior may vary across vLLM, Transformers, GPTQModel, and GPTQ-Marlin versions. ## Files Key files: - `model.safetensors.index.json` - `model-00001-of-00005.safetensors` through `model-00005-of-00005.safetensors` - `config.json` - `quantize_config.json` - `processor_config.json` - `tokenizer.json` - `UPLOAD_MANIFEST.json` `UPLOAD_MANIFEST.json` records the upload guardrail checks and artifact inspection summary. ## References - Source model: [`nex-agi/Nex-N2-mini`](https://huggingface.co/nex-agi/Nex-N2-mini) - GPTQ-Pro tooling: [`groxaxo/GPTQ-Pro`](https://github.com/groxaxo/GPTQ-Pro) - Reference recipe: [`groxaxo/Qwen3.6-27B-GPTQ-Pro-4bit`](https://huggingface.co/groxaxo/Qwen3.6-27B-GPTQ-Pro-4bit) - Terminal-Bench: [`laude-institute/terminal-bench`](https://github.com/laude-institute/terminal-bench) ## Individual Project Notice This repository is an individual research project. It is not affiliated with, sponsored by, or endorsed by any employer or organization.