whaoyang commited on
Commit
ff17b9d
·
verified ·
1 Parent(s): f168fcb

Add Python conversion script to README

Browse files
Files changed (1) hide show
  1. README.md +59 -1
README.md CHANGED
@@ -18,4 +18,62 @@ Compatible with RKLLM runtime version: 1.2.x
18
 
19
  Pretty much anything by these folks: [marty1885](https://github.com/marty1885) and [happyme531](https://huggingface.co/happyme531)
20
 
21
- Converted with instructions from [airockchip/rknn-llm #240](https://github.com/airockchip/rknn-llm/issues/240#issuecomment-2831806613)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
18
 
19
  Pretty much anything by these folks: [marty1885](https://github.com/marty1885) and [happyme531](https://huggingface.co/happyme531)
20
 
21
+ ## Conversion Python script
22
+
23
+ Based on instructions from [airockchip/rknn-llm #240](https://github.com/airockchip/rknn-llm/issues/240#issuecomment-2831806613)
24
+
25
+ ### `gemma-3-conversion.py`
26
+ ```
27
+ from rkllm.api import RKLLM
28
+ from transformers import Gemma3Processor, Gemma3ForConditionalGeneration
29
+ import safetensors
30
+ import torch
31
+
32
+ # Unsloth version
33
+ modelpath = 'unsloth/gemma-3-4b-it'
34
+
35
+ model = Gemma3ForConditionalGeneration.from_pretrained(modelpath, device_map='cpu', torch_dtype=torch.bfloat16).eval()
36
+ processor = Gemma3Processor.from_pretrained(modelpath, use_fast=True)
37
+
38
+ model.language_model.save_pretrained('llm')
39
+ processor.save_pretrained('llm')
40
+
41
+ del model
42
+ model = None
43
+ del processor
44
+ processor = None
45
+
46
+ modelpath = 'llm'
47
+ savepath = 'llm/gemma-3-4b-it-g128.rkllm'
48
+
49
+ llm = RKLLM()
50
+
51
+ ret = llm.load_huggingface(model=modelpath, device='cpu')
52
+ if ret != 0:
53
+ print('Load model failed!')
54
+ exit(ret)
55
+
56
+
57
+ ret = llm.build(
58
+ do_quantization=True,
59
+ optimization_level=0,
60
+ quantized_dtype='w8a8',
61
+ # hybrid ratio of 25% gives a good balance
62
+ hybrid_rate=0.25,
63
+ max_context=4096 * 4,
64
+ quantized_algorithm='normal',
65
+ target_platform='rk3588',
66
+ num_npu_core=3,
67
+ extra_qparams=None,
68
+ dataset=None
69
+ )
70
+ if ret != 0:
71
+ print('Build model failed!')
72
+ exit(ret)
73
+
74
+
75
+ ret = llm.export_rkllm(savepath)
76
+ if ret != 0:
77
+ print('Export model failed!')
78
+ exit(ret)
79
+ ```