BlackSamorez commited on
Commit
8be6cbf
·
verified ·
1 Parent(s): 1894e1e

Update README for swiss-ai/Apertus-v1.1-0.5B-Instruct-MLX-INT3

Browse files
Files changed (1) hide show
  1. README.md +7 -16
README.md CHANGED
@@ -54,7 +54,7 @@ Instead of standard pre-training, Apertus-v1.1 models were created using pre-tra
54
 
55
  This model family includes base pre-trained models and instruction-tuned models.
56
 
57
- For instruction-tuned models, we additionally provide high-quality quantization-aware distillation (QAD) checkpoints, obtained via the official [`qat-suite`](https://github.com/swiss-ai/qat-suite). We provide FP8 and NVFP4A16 checkpoints with vLLM inference in mind and INT3-6 checkpoints optimized for mobile usage on Apple devices.
58
 
59
  The full list of released checkpoints is shown below:
60
 
@@ -73,22 +73,18 @@ For more details refer to the original Apertus [technical report](https://arxiv.
73
 
74
  ## How to use
75
 
76
- The modeling code for Apertus is available in transformers `v4.56.0` and later, so make sure to upgrade your transformers version. You can also load the model with the latest `vLLM` which uses transformers as a backend.
77
  ```bash
78
- pip install -U transformers
79
  ```
80
 
81
  ```python
82
- from transformers import AutoModelForCausalLM, AutoTokenizer
83
 
84
  model_name = "swiss-ai/Apertus-v1.1-0.5B-Instruct-MLX-INT3"
85
- device = "cuda" # for GPU usage or "cpu" for CPU usage
86
 
87
- # load the tokenizer and the model
88
- tokenizer = AutoTokenizer.from_pretrained(model_name)
89
- model = AutoModelForCausalLM.from_pretrained(
90
- model_name,
91
- ).to(device)
92
 
93
  # prepare the model input
94
  prompt = "Give me a brief explanation of gravity in simple terms."
@@ -101,14 +97,9 @@ text = tokenizer.apply_chat_template(
101
  tokenize=False,
102
  add_generation_prompt=True,
103
  )
104
- model_inputs = tokenizer([text], return_tensors="pt", add_special_tokens=False).to(model.device)
105
 
106
  # Generate the output
107
- generated_ids = model.generate(**model_inputs, max_new_tokens=32768)
108
-
109
- # Get and decode the output
110
- output_ids = generated_ids[0][len(model_inputs.input_ids[0]) :]
111
- print(tokenizer.decode(output_ids, skip_special_tokens=True))
112
  ```
113
 
114
  >[!TIP]
 
54
 
55
  This model family includes base pre-trained models and instruction-tuned models.
56
 
57
+ For instruction-tuned models, we additionally provide high-quality quantization-aware distillation (QAD) checkpoints, obtained via the official [`qat-suite`]([https://github.com/swiss-ai/qat-suite](https://github.com/swiss-ai/qat-suite)). We provide FP8 and NVFP4A16 checkpoints with vLLM inference in mind and INT3-6 checkpoints optimized for mobile usage on Apple devices.
58
 
59
  The full list of released checkpoints is shown below:
60
 
 
73
 
74
  ## How to use
75
 
76
+ The modeling code for this Apertus MLX quantization is available via `mlx-lm`.
77
  ```bash
78
+ pip install mlx-lm
79
  ```
80
 
81
  ```python
82
+ from mlx_lm import load, generate
83
 
84
  model_name = "swiss-ai/Apertus-v1.1-0.5B-Instruct-MLX-INT3"
 
85
 
86
+ # load the model and tokenizer
87
+ model, tokenizer = load(model_name)
 
 
 
88
 
89
  # prepare the model input
90
  prompt = "Give me a brief explanation of gravity in simple terms."
 
97
  tokenize=False,
98
  add_generation_prompt=True,
99
  )
 
100
 
101
  # Generate the output
102
+ response = generate(model, tokenizer, prompt=text, verbose=True, max_tokens=32768)
 
 
 
 
103
  ```
104
 
105
  >[!TIP]