peculiar-ragdoll commited on
Commit
c147189
·
verified ·
1 Parent(s): 9e8327f

cards: document unrestricted reasoning/output budget defaults

Browse files
Files changed (1) hide show
  1. README.md +2 -0
README.md CHANGED
@@ -124,6 +124,8 @@ the next section, and measure before you rely on it.
124
 
125
  Sampling: `temperature 0.6`, `top_p 0.95`, `top_k 20`, `min_p 0` for agentic coding. For cybersecurity/CTF work, swap to `top_k 40`, `min_p 0.05` (same temperature and top_p), tested on the GGUF Q4 build. This is a looser configuration leading to more divergent and exploratory thinking, which leads to more solutions on Q4 but might create issues and non-convergence on lower quants.
126
 
 
 
127
  **Prefer to keep the files yourself?**
128
 
129
  ```bash
 
124
 
125
  Sampling: `temperature 0.6`, `top_p 0.95`, `top_k 20`, `min_p 0` for agentic coding. For cybersecurity/CTF work, swap to `top_k 40`, `min_p 0.05` (same temperature and top_p), tested on the GGUF Q4 build. This is a looser configuration leading to more divergent and exploratory thinking, which leads to more solutions on Q4 but might create issues and non-convergence on lower quants.
126
 
127
+ Budget: `mlx_vlm` has no unlimited default and requires an explicit `--max-tokens`; the `512` in the examples is sized for a one-shot demo prompt, not for real work. Give real work a generous ceiling — `32768` if you cap it at all. A low token budget degrades overall performance and will not necessarily make the model converge on the correct answer any faster. This model is much better than other 35B-A3B builds at spending fewer tokens and less time in total over the course of a problem — it knows when it needs to cook and when it is done — which makes high budgets, or no budget at all, both the safer and the better setting.
128
+
129
  **Prefer to keep the files yourself?**
130
 
131
  ```bash