Text Generation
Transformers
Safetensors
qwen3_5_moe
image-text-to-text
ornith
Mixture of Experts
gptq
gptq-pro
gptqmodel
foem
marlin
vllm
int4
quantized
long-context
tool-use
function-calling
terminal-bench
code
conversational
4-bit precision
Instructions to use XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256") model = AutoModelForMultimodalLM.from_pretrained("XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256
- SGLang
How to use XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256 with Docker Model Runner:
docker model run hf.co/XReyRobert/Ornith-1.0-35B-GPTQ-Pro-FOEM-4bit-g128-ns256
Align support buttons in model card
Browse files
README.md
CHANGED
|
@@ -33,12 +33,19 @@ datasets:
|
|
| 33 |
|
| 34 |
# Ornith-1.0-35B GPTQ-Pro FOEM 4-bit g128 ns256
|
| 35 |
|
| 36 |
-
<p>
|
| 37 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
<svg viewBox="0 0 24 24" fill="currentColor" aria-hidden="true" style="width: 16px; height: 16px; flex: 0 0 auto;"><path d="M18.244 2.25h3.308l-7.227 8.26 8.502 11.24H16.17l-5.214-6.817L4.99 21.75H1.68l7.73-8.835L1.254 2.25H8.08l4.713 6.231zm-1.161 17.52h1.833L7.084 4.126H5.117z"></path></svg>
|
| 39 |
Follow @xreyrobert
|
| 40 |
</a>
|
| 41 |
-
</
|
|
|
|
|
|
|
|
|
|
| 42 |
|
| 43 |
This is a GPTQ-Pro 4-bit quantization of
|
| 44 |
[`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B).
|
|
|
|
| 33 |
|
| 34 |
# Ornith-1.0-35B GPTQ-Pro FOEM 4-bit g128 ns256
|
| 35 |
|
| 36 |
+
<p style="margin: 10px 0 6px 0; color: #475569; font-size: 14px; line-height: 1.5;">
|
| 37 |
+
These models are built and maintained on rented GPU compute. If you want to show some appreciation, a follow on X or a coffee helps keep the releases coming.
|
| 38 |
+
</p>
|
| 39 |
+
|
| 40 |
+
<div style="display: flex; flex-wrap: wrap; align-items: center; gap: 10px; margin: 8px 0 18px 0;">
|
| 41 |
+
<a href="https://x.com/xreyrobert" target="_blank" rel="noopener noreferrer" class="follow-link" style="display: inline-flex; align-items: center; justify-content: center; gap: 8px; min-height: 36px; box-sizing: border-box; padding: 9px 14px; border: 1px solid #111827; border-radius: 999px; background: #111827; color: #ffffff; text-decoration: none; font-weight: 700; font-size: 14px; line-height: 1;">
|
| 42 |
<svg viewBox="0 0 24 24" fill="currentColor" aria-hidden="true" style="width: 16px; height: 16px; flex: 0 0 auto;"><path d="M18.244 2.25h3.308l-7.227 8.26 8.502 11.24H16.17l-5.214-6.817L4.99 21.75H1.68l7.73-8.835L1.254 2.25H8.08l4.713 6.231zm-1.161 17.52h1.833L7.084 4.126H5.117z"></path></svg>
|
| 43 |
Follow @xreyrobert
|
| 44 |
</a>
|
| 45 |
+
<a href="https://buymeacoffee.com/xrrxrr" target="_blank" rel="noopener noreferrer" class="support-link" style="display: inline-flex; align-items: center; justify-content: center; gap: 8px; min-height: 36px; box-sizing: border-box; padding: 9px 14px; border: 1px solid #f59e0b; border-radius: 999px; background: #fbbf24; color: #111827; text-decoration: none; font-weight: 700; font-size: 14px; line-height: 1;">
|
| 46 |
+
Buy me a coffee
|
| 47 |
+
</a>
|
| 48 |
+
</div>
|
| 49 |
|
| 50 |
This is a GPTQ-Pro 4-bit quantization of
|
| 51 |
[`deepreinforce-ai/Ornith-1.0-35B`](https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B).
|