jedisct1 commited on
Commit
639fec6
·
verified ·
1 Parent(s): 0ad1e17

Refresh model card

Browse files
Files changed (1) hide show
  1. README.md +43 -9
README.md CHANGED
@@ -14,17 +14,51 @@ tags:
14
 
15
  # Qwen3.6-35B-go-v2 4-bit MLX
16
 
17
- This is the MLX 4-bit package of the Go-v2 fine-tuned Qwen3.6-35B-A3B model without native MTP tensors.
18
 
19
- It is intended as the plain compatibility sibling for loaders that do not currently accept Qwen3.5 MoE native-MTP weights. The native-MTP sibling is [jedisct1/Qwen3.6-35B-go-v2-MTP-4bit.mlx](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-MTP-4bit.mlx).
20
 
21
- ## Compatibility
22
 
23
- - Native MTP tensors are intentionally removed.
24
- - The repository name does not include `-MTP`.
25
- - This package keeps the fine-tuned main model weights, tokenizer files, generation config, and chat template.
26
- - There are no root LoRA adapter files; this is a standalone MLX artifact.
27
 
28
- ## Source
 
 
 
29
 
30
- The artifact was derived from the repaired Go-v2 MLX package by removing `mtp.safetensors`, deleting MTP weight-map entries, and clearing the MTP config fields.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14
 
15
  # Qwen3.6-35B-go-v2 4-bit MLX
16
 
17
+ A Go-focused Qwen3.6-35B-A3B model for Apple Silicon, packaged in MLX.
18
 
19
+ Use it as a coding assistant for Go projects: generating focused patches, explaining diffs, tightening tests, reading tool outputs, and making small repo-aware edits. It was tested with [Swival](https://swival.dev) on file-editing and command-running workflows.
20
 
21
+ This is the plain 4-bit compatibility variant. It does not include native MTP tensors, so it is the best starting point if your MLX loader does not support MTP sidecars.
22
 
23
+ ## Which Variant Should I Use?
 
 
 
24
 
25
+ - **Use this repo** if you want the smallest plain MLX package or need a loader-friendly non-MTP model.
26
+ - Use [`jedisct1/Qwen3.6-35B-go-v2-8bit.mlx`](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-8bit.mlx) if you want more precision without native MTP.
27
+ - Use [`jedisct1/Qwen3.6-35B-go-v2-bf16.mlx`](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-bf16.mlx) if you want full precision without native MTP.
28
+ - Use [`jedisct1/Qwen3.6-35B-go-v2-MTP-4bit.mlx`](https://huggingface.co/jedisct1/Qwen3.6-35B-go-v2-MTP-4bit.mlx) if your runtime supports native MTP and you want the faster MTP path.
29
 
30
+ ## Usage
31
+
32
+ Requires [mlx-lm](https://github.com/ml-explore/mlx-examples/tree/main/llms/mlx_lm):
33
+
34
+ ```bash
35
+ pip install mlx-lm
36
+ ```
37
+
38
+ ```python
39
+ from mlx_lm import load, generate
40
+
41
+ model, tokenizer = load("jedisct1/Qwen3.6-35B-go-v2-4bit.mlx")
42
+
43
+ messages = [
44
+ {"role": "system", "content": "You are an expert Go developer."},
45
+ {"role": "user", "content": "Generate a focused patch that replaces the manual retry loop in fetchUser() with the shared retry helper."},
46
+ ]
47
+ prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
48
+ response = generate(model, tokenizer, prompt=prompt, max_tokens=500)
49
+ print(response)
50
+ ```
51
+
52
+ ## What It Is Good At
53
+
54
+ - Writing idiomatic Go patches from a concise change request.
55
+ - Explaining Go diffs in commit-message style.
56
+ - Following tool-calling workflows where it needs to inspect files before editing.
57
+ - Keeping changes focused instead of turning small fixes into broad rewrites.
58
+ - Working with tests, command output, and repository context.
59
+
60
+ ## Limitations
61
+
62
+ - Outputs should be reviewed before use, especially patches that touch production systems.
63
+ - The model works best on focused Go changes, tests, and explanations. Very large refactors may need to be split into smaller steps.
64
+ - Tool calling depends on the runtime and client preserving the chat template and tool schema format.