MisterOss commited on
Commit
9f4d156
·
verified ·
1 Parent(s): eae5ed4

Clarify sharding restrictions and data-parallel support

Browse files
Files changed (1) hide show
  1. README.md +6 -2
README.md CHANGED
@@ -78,8 +78,8 @@ tokenizer = AutoTokenizer.from_pretrained(
78
  model = AutoModelForCausalLM.from_pretrained(
79
  model_id,
80
  trust_remote_code=True,
81
- torch_dtype="auto",
82
- device_map="auto",
83
  attn_implementation="sdpa",
84
  )
85
 
@@ -101,6 +101,10 @@ answer = tokenizer.decode(
101
  print(answer)
102
  ```
103
 
 
 
 
 
104
  ### Supported Transformers inference paths
105
 
106
  | Attention backend | Cache | Status |
 
78
  model = AutoModelForCausalLM.from_pretrained(
79
  model_id,
80
  trust_remote_code=True,
81
+ dtype="auto",
82
+ device_map="cuda",
83
  attn_implementation="sdpa",
84
  )
85
 
 
101
  print(answer)
102
  ```
103
 
104
+ The Transformers implementation supports one complete model replica per process. Automatic or explicit model sharding, tensor parallelism, and pipeline parallelism are not currently supported.
105
+
106
+ For offline multi-GPU inference, independent replicas can be launched with one process per GPU using `device_map={"": local_rank}`. For production serving, continuous batching, and multi-GPU request scheduling, use the [Limite vLLM integration](https://github.com/paradigma-inc/limite-violetto).
107
+
108
  ### Supported Transformers inference paths
109
 
110
  | Attention backend | Cache | Status |