FreedomAISVR commited on
Commit
ccd2374
Β·
verified Β·
1 Parent(s): d324434

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +15 -20
README.md CHANGED
@@ -19,9 +19,7 @@ base_model: Qwen/Qwen-AgentWorld-35B-A3B
19
 
20
  # Qwen AgentWorld 35B-A3B β€” NVFP4 GGUF
21
 
22
- NVFP4 quantization of [Qwen/Qwen-AgentWorld-35B-A3B](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B), a 35B parameter Mixture-of-Experts model with 3B active parameters, designed for agent tasks and world modeling .
23
-
24
- > **Note: This model has NO vision support.** The official Qwen/Qwen-AgentWorld-35B-A3B checkpoint ships with \language_model_only: True\ β€” the vision encoder weights are not included. Despite the \image-text-to-text\ tag on the source model and the vision config in config.json, this GGUF is text-only. Do not attempt to send images to this model.
25
 
26
  ## About the Model
27
 
@@ -29,14 +27,16 @@ Qwen AgentWorld is a specialized variant of the Qwen 3.5 MoE architecture optimi
29
 
30
  - **Agent tasks** β€” tool calling, function execution, environment simulation
31
  - **World modeling** β€” understanding and predicting environment states
32
- **35B total parameters** with only **3B active per token** (256 experts, 8 active)
 
33
  - **Efficient inference** β€” MoE architecture activates only a fraction of parameters
34
 
35
  ## Architecture
36
 
37
  - **Text model**: Qwen3.5 MoE β€” 40 layers, 2048 hidden, 256 experts (8 active/token)
38
- - **Vision encoder**: 27-layer SigLIP-style, 1152 hidden, patch_size 16
39
  - **Vocabulary**: 248,320 tokens
 
40
  ## Quantization
41
 
42
  This GGUF was quantized from the BF16 safetensors using [llama.cpp](https://github.com/ggerganov/llama.cpp) (build 537). The source weights were converted to F16 GGUF, then quantized to NVFP4 format.
@@ -47,39 +47,34 @@ NVFP4 (NVIDIA FP4) uses 4-bit floating point quantization optimized for NVIDIA B
47
 
48
  | File | Size | Description |
49
  |------|------|-------------|
50
- | `qwen-agentworld-35b-a3b-nvfp4.gguf` | ~18.4 GB | NVFP4 quantized model weights
51
- | `mmproj-qwen-agentworld-35b-a3b-f16.gguf` | ~0.84 GB | Vision projector (BF16) |
52
 
53
  ## Usage
54
 
55
  ### llama.cpp
56
 
57
- ```bash
58
- # Server mode with OpenAI-compatible API
59
  llama-server \
60
  -m qwen-agentworld-35b-a3b-nvfp4.gguf \
 
61
  -ngl 99 \
62
  --host 0.0.0.0 \
63
  --port 8080
64
-
65
- # Direct inference
66
- llama-cli \
67
- -m qwen-agentworld-35b-a3b-nvfp4.gguf \
68
- -ngl 99 \
69
- -p "Analyze this image and describe what you see"
70
- ```
71
 
72
  ### LM Studio
73
 
74
- 1. Download the GGUF file from this repository
75
- 2. Load the GGUF file in LM Studio (vision is embedded, no mmproj needed)
76
- 3. Set GPU offload layers to maximum
 
77
 
78
  ## Hardware Requirements
79
 
80
  - **Minimum**: 20 GB VRAM for partial offload
81
  - **Recommended**: 24+ GB VRAM for full GPU offload
82
- - **Disk**: ~18.4 GB
83
 
84
  ## License
85
 
 
19
 
20
  # Qwen AgentWorld 35B-A3B β€” NVFP4 GGUF
21
 
22
+ NVFP4 quantization of [Qwen/Qwen-AgentWorld-35B-A3B](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B), a 35B parameter Mixture-of-Experts model with 3B active parameters, designed for agent tasks and world modeling with vision support.
 
 
23
 
24
  ## About the Model
25
 
 
27
 
28
  - **Agent tasks** β€” tool calling, function execution, environment simulation
29
  - **World modeling** β€” understanding and predicting environment states
30
+ - **Vision understanding** β€” multimodal image input via separate mmproj vision projector
31
+ - **35B total parameters** with only **3B active per token** (256 experts, 8 active)
32
  - **Efficient inference** β€” MoE architecture activates only a fraction of parameters
33
 
34
  ## Architecture
35
 
36
  - **Text model**: Qwen3.5 MoE β€” 40 layers, 2048 hidden, 256 experts (8 active/token)
37
+ - **Vision encoder**: 27-layer SigLIP-style, 1152 hidden, patch_size 16 (via mmproj)
38
  - **Vocabulary**: 248,320 tokens
39
+
40
  ## Quantization
41
 
42
  This GGUF was quantized from the BF16 safetensors using [llama.cpp](https://github.com/ggerganov/llama.cpp) (build 537). The source weights were converted to F16 GGUF, then quantized to NVFP4 format.
 
47
 
48
  | File | Size | Description |
49
  |------|------|-------------|
50
+ | qwen-agentworld-35b-a3b-nvfp4.gguf | ~18.4 GB | NVFP4 quantized model weights |
51
+ | mmproj-qwen-agentworld-35b-a3b-f16.gguf | ~843 MB | Vision projector (BF16) |
52
 
53
  ## Usage
54
 
55
  ### llama.cpp
56
 
57
+ `ash
 
58
  llama-server \
59
  -m qwen-agentworld-35b-a3b-nvfp4.gguf \
60
+ --mmproj mmproj-qwen-agentworld-35b-a3b-f16.gguf \
61
  -ngl 99 \
62
  --host 0.0.0.0 \
63
  --port 8080
64
+ `
 
 
 
 
 
 
65
 
66
  ### LM Studio
67
 
68
+ 1. Download both files from this repository
69
+ 2. Load the main GGUF file in LM Studio
70
+ 3. Load the mmproj file for vision support
71
+ 4. Set GPU offload layers to maximum
72
 
73
  ## Hardware Requirements
74
 
75
  - **Minimum**: 20 GB VRAM for partial offload
76
  - **Recommended**: 24+ GB VRAM for full GPU offload
77
+ - **Disk**: ~19.2 GB
78
 
79
  ## License
80