Jianshu001 commited on
Commit
3b453d0
·
verified ·
1 Parent(s): 80066c6

Update model card for the mlx-community namespace

Browse files
Files changed (1) hide show
  1. README.md +5 -5
README.md CHANGED
@@ -55,7 +55,7 @@ Smoke-tested after conversion with released mlx-lm 0.31.3: `17 * 23` → `391` a
55
 
56
  WikiText-2 test perplexity (128 × 512 tokens, lower is better) and generation speed on an M4 Max 128GB, all measured the same way:
57
 
58
- | | [bf16](https://huggingface.co/Jianshu001/K2-Horizon-3.7B-bf16) | [8-bit](https://huggingface.co/Jianshu001/K2-Horizon-3.7B-8bit) | [4-bit](https://huggingface.co/Jianshu001/K2-Horizon-3.7B-4bit) |
59
  |---|---|---|---|
60
  | Bits/weight | 16 | 8.5 | 6.5 |
61
  | Disk | 10.1 GB | 5.4 GB | 4.1 GB |
@@ -63,19 +63,19 @@ WikiText-2 test perplexity (128 × 512 tokens, lower is better) and generation s
63
  | WikiText-2 perplexity | 17.552 | 17.544 (-0.05%) | 18.324 (+4.4%) |
64
  | Generation | 50.3 tok/s | 84.2 tok/s | 109.6 tok/s |
65
 
66
- Perplexity is a coarse signal. Test the versions on your own workload before picking one. Other K2-Horizon sizes: the [K2-Horizon Dense collection](https://huggingface.co/collections/Jianshu001/k2-horizon-dense-6abb69400f675cbe43b29f82).
67
 
68
  ## Usage
69
 
70
  ```bash
71
- mlx_lm.generate --model Jianshu001/K2-Horizon-3.7B-4bit --trust-remote-code --prompt "Explain mixture-of-experts in two sentences." --max-tokens 2048
72
  ```
73
 
74
  ```python
75
  from mlx_lm import load, generate
76
 
77
- # Newer mlx-lm versions need trust_remote_code=True; on mlx-lm 0.31.3 use load("Jianshu001/K2-Horizon-3.7B-4bit").
78
- model, tokenizer = load("Jianshu001/K2-Horizon-3.7B-4bit", trust_remote_code=True)
79
  messages = [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}]
80
  prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
81
  print(generate(model, tokenizer, prompt, max_tokens=2048))
 
55
 
56
  WikiText-2 test perplexity (128 × 512 tokens, lower is better) and generation speed on an M4 Max 128GB, all measured the same way:
57
 
58
+ | | [bf16](https://huggingface.co/mlx-community/K2-Horizon-3.7B-bf16) | [8-bit](https://huggingface.co/mlx-community/K2-Horizon-3.7B-8bit) | [4-bit](https://huggingface.co/mlx-community/K2-Horizon-3.7B-4bit) |
59
  |---|---|---|---|
60
  | Bits/weight | 16 | 8.5 | 6.5 |
61
  | Disk | 10.1 GB | 5.4 GB | 4.1 GB |
 
63
  | WikiText-2 perplexity | 17.552 | 17.544 (-0.05%) | 18.324 (+4.4%) |
64
  | Generation | 50.3 tok/s | 84.2 tok/s | 109.6 tok/s |
65
 
66
+ Perplexity is a coarse signal. Test the versions on your own workload before picking one. Other K2-Horizon sizes: the [K2-Horizon collection](https://huggingface.co/collections/mlx-community/k2-horizon).
67
 
68
  ## Usage
69
 
70
  ```bash
71
+ mlx_lm.generate --model mlx-community/K2-Horizon-3.7B-4bit --trust-remote-code --prompt "Explain mixture-of-experts in two sentences." --max-tokens 2048
72
  ```
73
 
74
  ```python
75
  from mlx_lm import load, generate
76
 
77
+ # Newer mlx-lm versions need trust_remote_code=True; on mlx-lm 0.31.3 use load("mlx-community/K2-Horizon-3.7B-4bit").
78
+ model, tokenizer = load("mlx-community/K2-Horizon-3.7B-4bit", trust_remote_code=True)
79
  messages = [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}]
80
  prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
81
  print(generate(model, tokenizer, prompt, max_tokens=2048))