Clarify sharding restrictions and data-parallel support
Browse files
README.md
CHANGED
|
@@ -78,8 +78,8 @@ tokenizer = AutoTokenizer.from_pretrained(
|
|
| 78 |
model = AutoModelForCausalLM.from_pretrained(
|
| 79 |
model_id,
|
| 80 |
trust_remote_code=True,
|
| 81 |
-
|
| 82 |
-
device_map="
|
| 83 |
attn_implementation="sdpa",
|
| 84 |
)
|
| 85 |
|
|
@@ -101,6 +101,10 @@ answer = tokenizer.decode(
|
|
| 101 |
print(answer)
|
| 102 |
```
|
| 103 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 104 |
### Supported Transformers inference paths
|
| 105 |
|
| 106 |
| Attention backend | Cache | Status |
|
|
|
|
| 78 |
model = AutoModelForCausalLM.from_pretrained(
|
| 79 |
model_id,
|
| 80 |
trust_remote_code=True,
|
| 81 |
+
dtype="auto",
|
| 82 |
+
device_map="cuda",
|
| 83 |
attn_implementation="sdpa",
|
| 84 |
)
|
| 85 |
|
|
|
|
| 101 |
print(answer)
|
| 102 |
```
|
| 103 |
|
| 104 |
+
The Transformers implementation supports one complete model replica per process. Automatic or explicit model sharding, tensor parallelism, and pipeline parallelism are not currently supported.
|
| 105 |
+
|
| 106 |
+
For offline multi-GPU inference, independent replicas can be launched with one process per GPU using `device_map={"": local_rank}`. For production serving, continuous batching, and multi-GPU request scheduling, use the [Limite vLLM integration](https://github.com/paradigma-inc/limite-violetto).
|
| 107 |
+
|
| 108 |
### Supported Transformers inference paths
|
| 109 |
|
| 110 |
| Attention backend | Cache | Status |
|