sunil-pathak commited on
Commit
0d0ba7b
·
verified ·
1 Parent(s): 8144785

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +6 -4
README.md CHANGED
@@ -16,9 +16,10 @@ pipeline_tag: text-generation
16
 
17
  ## 📊 Performance Metrics
18
 
 
19
  - **Size:** 3.47 GB
20
- - **Speed (Generation):** 5.22 tokens/sec
21
- - **Speed (Prompt):** 10.85 tokens/sec
22
  - **KV Cache Usage:** 0.0143 GB
23
  - **Quantization:** Q6_K
24
 
@@ -31,7 +32,7 @@ This repository contains a **GGUF quantized version** of:
31
  - **Base Model:** gemma-3n-E2B-it
32
  - **Format:** GGUF (optimized for llama.cpp inference)
33
  - **Precision:** Q6_K
34
- - **Efficiency Score:** 1.5055 (TPS/GB)
35
 
36
  GGUF format provides:
37
  - Fast loading via memory mapping
@@ -57,7 +58,8 @@ GGUF format provides:
57
  | Format | GGUF |
58
  | Precision | Q6_K |
59
  | Runtime | llama.cpp |
60
- | Context Latency | 28.58s |
 
61
  | Memory (KV) | 0.0143 GB |
62
 
63
  ---
 
16
 
17
  ## 📊 Performance Metrics
18
 
19
+ - **Hardware:** Intel(R) Xeon(R) CPU @ 2.20GHz (4 vCPUs)
20
  - **Size:** 3.47 GB
21
+ - **Speed (Generation):** 5.41 tokens/sec
22
+ - **Speed (Prompt):** 10.76 tokens/sec
23
  - **KV Cache Usage:** 0.0143 GB
24
  - **Quantization:** Q6_K
25
 
 
32
  - **Base Model:** gemma-3n-E2B-it
33
  - **Format:** GGUF (optimized for llama.cpp inference)
34
  - **Precision:** Q6_K
35
+ - **Efficiency Score:** 1.5603 (TPS/GB)
36
 
37
  GGUF format provides:
38
  - Fast loading via memory mapping
 
58
  | Format | GGUF |
59
  | Precision | Q6_K |
60
  | Runtime | llama.cpp |
61
+ | Benchmark Hardware | Intel(R) Xeon(R) CPU @ 2.20GHz (4 vCPUs) |
62
+ | Context Latency | 29.69s |
63
  | Memory (KV) | 0.0143 GB |
64
 
65
  ---