Instructions to use jedisct1/Qwen3.6-35B-go-v2-4bit.mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use jedisct1/Qwen3.6-35B-go-v2-4bit.mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download jedisct1/Qwen3.6-35B-go-v2-4bit.mlx --local-dir Qwen3.6-35B-go-v2-4bit.mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Refresh model card
Browse files
README.md
CHANGED
|
@@ -14,17 +14,51 @@ tags:
|
|
| 14 |
|
| 15 |
# Qwen3.6-35B-go-v2 4-bit MLX
|
| 16 |
|
| 17 |
-
|
| 18 |
|
| 19 |
-
|
| 20 |
|
| 21 |
-
|
| 22 |
|
| 23 |
-
|
| 24 |
-
- The repository name does not include `-MTP`.
|
| 25 |
-
- This package keeps the fine-tuned main model weights, tokenizer files, generation config, and chat template.
|
| 26 |
-
- There are no root LoRA adapter files; this is a standalone MLX artifact.
|
| 27 |
|
| 28 |
-
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
# Qwen3.6-35B-go-v2 4-bit MLX
|
| 16 |
|
| 17 |
+
A Go-focused Qwen3.6-35B-A3B model for Apple Silicon, packaged in MLX.
|
| 18 |
|
| 19 |
+
Use it as a coding assistant for Go projects: generating focused patches, explaining diffs, tightening tests, reading tool outputs, and making small repo-aware edits. It was tested with [Swival](https://swival.dev) on file-editing and command-running workflows.
|
| 20 |
|
| 21 |
+
This is the plain 4-bit compatibility variant. It does not include native MTP tensors, so it is the best starting point if your MLX loader does not support MTP sidecars.
|
| 22 |
|
| 23 |
+
## Which Variant Should I Use?
|
|
|
|
|
|
|
|
|
|
| 24 |
|
| 25 |
+
- **Use this repo** if you want the smallest plain MLX package or need a loader-friendly non-MTP model.
|
| 26 |
+
- Use [`jedisct1/Qwen3.6-35B-go-v2-8bit.mlx`](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-8bit.mlx) if you want more precision without native MTP.
|
| 27 |
+
- Use [`jedisct1/Qwen3.6-35B-go-v2-bf16.mlx`](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-bf16.mlx) if you want full precision without native MTP.
|
| 28 |
+
- Use [`jedisct1/Qwen3.6-35B-go-v2-MTP-4bit.mlx`](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-MTP-4bit.mlx) if your runtime supports native MTP and you want the faster MTP path.
|
| 29 |
|
| 30 |
+
## Usage
|
| 31 |
+
|
| 32 |
+
Requires [mlx-lm](https://github.com/ml-explore/mlx-examples/tree/main/llms/mlx_lm):
|
| 33 |
+
|
| 34 |
+
```bash
|
| 35 |
+
pip install mlx-lm
|
| 36 |
+
```
|
| 37 |
+
|
| 38 |
+
```python
|
| 39 |
+
from mlx_lm import load, generate
|
| 40 |
+
|
| 41 |
+
model, tokenizer = load("jedisct1/Qwen3.6-35B-go-v2-4bit.mlx")
|
| 42 |
+
|
| 43 |
+
messages = [
|
| 44 |
+
{"role": "system", "content": "You are an expert Go developer."},
|
| 45 |
+
{"role": "user", "content": "Generate a focused patch that replaces the manual retry loop in fetchUser() with the shared retry helper."},
|
| 46 |
+
]
|
| 47 |
+
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
| 48 |
+
response = generate(model, tokenizer, prompt=prompt, max_tokens=500)
|
| 49 |
+
print(response)
|
| 50 |
+
```
|
| 51 |
+
|
| 52 |
+
## What It Is Good At
|
| 53 |
+
|
| 54 |
+
- Writing idiomatic Go patches from a concise change request.
|
| 55 |
+
- Explaining Go diffs in commit-message style.
|
| 56 |
+
- Following tool-calling workflows where it needs to inspect files before editing.
|
| 57 |
+
- Keeping changes focused instead of turning small fixes into broad rewrites.
|
| 58 |
+
- Working with tests, command output, and repository context.
|
| 59 |
+
|
| 60 |
+
## Limitations
|
| 61 |
+
|
| 62 |
+
- Outputs should be reviewed before use, especially patches that touch production systems.
|
| 63 |
+
- The model works best on focused Go changes, tests, and explanations. Very large refactors may need to be split into smaller steps.
|
| 64 |
+
- Tool calling depends on the runtime and client preserving the chat template and tool schema format.
|