inferencerlabs commited on
Commit
cc8020e
·
verified ·
1 Parent(s): 6681050

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -11,7 +11,7 @@ pipeline_tag: text-generation
11
  **See NVIDIA-Nemotron-3-Super-120B-A12B MLX in action - [demonstration video](https://youtu.be/MzeRCbnOg9Q)**
12
 
13
  #### Tested on a M3 Ultra 512GB RAM using [Inferencer app v1.10.6](https://inferencer.com)
14
- - Single inference ~41.5 tokens/s @ 1000 tokens (in debug mode - release figured coming soon)
15
  - Batched inference ~ total tokens/s across five inferences
16
  - Memory usage: ~67 GiB
17
 
 
11
  **See NVIDIA-Nemotron-3-Super-120B-A12B MLX in action - [demonstration video](https://youtu.be/MzeRCbnOg9Q)**
12
 
13
  #### Tested on a M3 Ultra 512GB RAM using [Inferencer app v1.10.6](https://inferencer.com)
14
+ - Single inference ~49.6 tokens/s @ 1000 tokens
15
  - Batched inference ~ total tokens/s across five inferences
16
  - Memory usage: ~67 GiB
17