JoyFox Qwen3.6 35B-A3B RP Aggressive — NVFP4 FTW

A text-only, FreeToken-ready FTW conversion of joyfox/JoyFox-Qwen3.6-35B-A3B-RP-Aggressive.

This is a quantized derivative, not a new finetune. The goal was to preserve the behavior of the original BF16 JoyFox model as closely as practical while reducing its storage and runtime footprint for local inference.

Highlights

  • FreeToken FTW format
  • Qwen3.6 35B-A3B MoE
  • Text-only
  • Vision removed
  • MTP removed
  • Mixed NVFP4 + FP8 + BF16/FP32 precision
  • Approximately 19.456 GiB of model tensors
  • Approximately 19.475 GiB for the complete self-contained model folder
  • Tested successfully with FreeToken 0.1.2+g141c31a8d
  • Validated with arithmetic, constrained-format, explanatory, and roleplay prompts

The original BF16 source contained approximately 66.965 GiB of model tensors, so this conversion reduces the working model by roughly 47.5 GiB.


Quantization recipe

This build uses a mixed-precision ModelOpt-style layout supported by FreeToken's Qwen3.5/Qwen3.6 implementation.

Model component Final format Reason
Routed MoE experts Native NVFP4 E2M1 Largest part of the model; provides most of the size reduction
Shared expert Native NVFP4 E2M1 Supported by FreeToken's native W4A16 path
Language head Native NVFP4 E2M1 Reduces size substantially while using FreeToken's dedicated NVFP4 language-head implementation
Full-attention projections FP8 E4M3 Preserves more precision than NVFP4
Gated DeltaNet projection weights FP8 E4M3 Supported by FreeToken's W8A16 path
Input embedding BF16 Preserved because embeddings are sensitive to quantization
Gated DeltaNet in_proj_a / in_proj_b BF16 Preserved for compatibility with FreeToken's implementation
Norms and routers Original precision Small and structurally sensitive
Mamba / Gated DeltaNet state parameters BF16 / FP32 Small enough to preserve at source precision
Vision Removed Not required for this text-only build
MTP Removed Not required for this build

FreeToken's required Qwen3.6 Gemma-style norm transformation was applied during FTW construction.


Tensor accounting

Every source tensor was individually classified and accounted for.

  • 31,685 source tensors
  • 31,685 accounted-for tensors
  • 30,841 converted to NVFP4
  • 130 converted to FP8
  • 362 retained or fused at source precision
  • 352 removed
  • 0 missing source entries
  • 0 duplicate source entries

Removed tensors

  • 333 vision tensors
    • 893,142,496 bytes
    • approximately 0.832 GiB
  • 19 MTP tensors
    • 1,689,281,536 bytes
    • approximately 1.573 GiB

Final FTW structure

  • 975 stored tensors
  • 735 normal model tensors
  • 240 native NVFP4 expert-bank tensors
  • 40 complete expert layers
  • 3 FTW shards
  • No vision tensors
  • No MTP tensors
  • No missing expert banks
  • No overlapping tensors
  • No out-of-bounds tensor entries

Quantization error

Measured relative RMS error during conversion:

Format Mean relative RMS error Maximum relative RMS error
FP8 2.650% 2.660%
NVFP4 9.487% 9.543%

The NVFP4 error distribution remained consistent across the routed expert matrices, shared experts, and language head. No anomalous tensors exceeded the validation threshold used during conversion.


FreeToken compatibility

Tested with:

FreeToken 0.1.2+g141c31a8d

FreeToken detected the intended layout:

  • Routed experts: nvfp4
  • Dense/shared expert: nvfp4
  • Attention: fp8_pertensor
  • Language head: nvfp4
  • MoE: enabled across 40 layers

The checkpoint successfully:

  • Loaded through the FTW fast path
  • Loaded the native NVFP4 expert banks
  • Created an NVFP4 expert cache
  • Allocated KV cache
  • Captured its CUDA graph
  • Completed prefill warmup
  • Reached ready-to-serve state
  • Produced coherent API responses

Validation

The final build was tested with real generation rather than structural checks alone.

Non-thinking tests included:

  1. Arithmetic: returned exactly 703
  2. A sky-color explanation: returned a correct answer in exactly two sentences
  3. Exact-format instruction: returned exactly Calm blue ocean
  4. Roleplay: returned coherent, in-character dialogue

Thinking mode was also coherent and independently calculated the arithmetic result correctly, but JoyFox can be unusually verbose while reasoning.

For ordinary chat and roleplay, non-thinking mode is recommended.

Example request setting:

"chat_template_kwargs": {
  "enable_thinking": false
}

Intended use

This build is primarily intended for:

  • Local roleplay
  • Creative writing
  • General chat
  • Long-form conversation
  • Coding and general Qwen3.6 tasks
  • FreeToken-based local inference on systems where BF16 is impractical

The original JoyFox model is specifically tuned toward roleplay and creative writing. This repository preserves that model behavior while changing its storage/inference representation.


Limitations

  • This is an FTW / FreeToken-oriented checkpoint, not a standard drop-in Transformers BF16 model.
  • Vision capability has been intentionally removed.
  • MTP has been intentionally removed.
  • Quantization can introduce small behavioral differences relative to the BF16 source.
  • Thinking mode can be very verbose with this model/template.
  • Tool-calling capability was not comprehensively benchmarked as part of this conversion.

Source model and license

Source:

joyfox/JoyFox-Qwen3.6-35B-A3B-RP-Aggressive

This derivative retains the source model's Apache License 2.0.

No additional training was performed for this release. Only quantization, text-only extraction, FTW conversion, and validation were performed.


Conversion summary

The conversion reduced the original 66.965 GiB BF16 model tensor footprint to approximately 19.456 GiB of FTW model tensors while preserving sensitive components at BF16/FP32 and using FreeToken-native NVFP4/FP8 paths for the large compute-heavy tensors.

The result is a compact, text-only JoyFox build intended to retain as much of the original BF16 model's RP and general-purpose behavior as practical while being substantially easier to run locally.

Downloads last month
304
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hoisd/JoyFox-Qwen3.6-35B-A3B-RP-Aggressive-NVFP4-ftw

Quantized
(4)
this model