Paiton Ornith 1.5 35B-A3B on AMD Radeon AI PRO R9700

Run Capicua25x/Ornith-1.5-35B-A3B-MXFP4-Quark-RDNA4 locally through Paiton's compiled AMD GPU runtime and an OpenAI-compatible API.

This repository contains compiled runtime artifacts, configuration and manifests only, about 4.46 MB. Model weights are downloaded directly from the original publisher and remain in your local cache. The .so files contain executable model graphs and GPU kernels. The private Paiton compiler is not required to use them and is not included.

Release: v1.0.0 路 GPU: Radeon AI PRO R9700 (gfx1201, 32 GB) 路 Batch: 1 路 Tensor parallelism: 1 路 Context: 8,192 tokens 路 Text only

Quick start

You need Linux x86-64, Docker and working ROCm device access through /dev/kfd and /dev/dri. The image includes the qualified ROCm 7.14, PyTorch, vLLM and Paiton runtime. This release is qualified on the R9700; other GPUs and runtime versions are not covered by that qualification.

docker run -d \
  --name paiton-ornith \
  --device /dev/kfd \
  --device /dev/dri \
  --group-add video \
  --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v paiton-ornith-cache:/models/cache \
  ghcr.io/eliovp/paiton-vllm-plugin:ornith15-mxfp4-rdna4-v1.0.0

The first start downloads the 22.9 GB target checkpoint and the 772 MB DFlash draft directly from the original Hugging Face repository at the revision below. Paiton losslessly splits the target checkpoint into smaller shards locally. Allow about 48 GB of free disk space for staging. The named volume preserves the downloads and staged model for later runs.

The container already bundles the same compiled artifacts published here. You do not need to download the .so separately for this quick start. The upstream repositories are public; no Hugging Face token is required. Run one example server at a time because both use port 8000.

Follow startup:

docker logs -f paiton-ornith

Wait for Application startup complete. Then, from another terminal:

curl --fail http://127.0.0.1:8000/v1/models
docker exec -it paiton-ornith paiton-chat --model ornith

Use /reset to clear the chat and /quit to leave it. Or call the API:

curl --fail http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"ornith","messages":[{"role":"user","content":"Explain why the sky is blue."}],"temperature":0,"max_tokens":256,"chat_template_kwargs":{"enable_thinking":false}}'

Model and execution details

The Hub quantization relationship identifies this compiled runtime package with its upstream quantized checkpoint. It does not indicate that this repository publishes a new set of quantized weights.

  • Original repository: Capicua25x/Ornith-1.5-35B-A3B-MXFP4-Quark-RDNA4.
  • Pinned weight revision: 9e488f46c0f7969f84c9923ee0256311cd50316e.
  • Execution format: Checkpoint-native Quark MXFP4; DFlash with up to 16 speculative tokens.
  • One active request; additional requests queue. Changing server flags does not add multi-batch or multi-GPU support to this release.

DFlash is enabled by default and uses the upstream repository's dflash-draft/ files at the same pinned revision as the target. Neither target nor draft weights are mirrored here. Add -e PAITON_ORNITH_DFLASH=0 to the Docker command to disable speculation. The released launcher still stages the draft files when preparing the model.

Published performance

The published R9700 benchmark reports 44.628 output tokens/s for Paiton with DFlash 16, averaged over two runs, versus 35.132 output tokens/s for the fastest obtained stock vLLM configuration without DFlash: 27.03% higher output throughput. This is a whole-package comparison with different speculation settings, not an isolated kernel speedup.

Each run used concurrency 1, three warmups and 12 measured requests with 256 requested input tokens and exactly 256 output tokens. The same target checkpoint, tokenizer and GPU were used. All 24 Paiton requests completed. The reported quality check was a deterministic natural-language sanity test; it does not establish broad task-quality equivalence.

See the published benchmark settings, results and reproduction command. These are existing release results; no new benchmark was run for this Hub package.

Download and use these Hub artifacts

For manual artifact use, install the Hugging Face CLI in your local Python environment, then download the release into a separate directory:

python3 -m pip install huggingface_hub
hf download EliovpAI/Ornith-1.5-35B-A3B-MXFP4-Paiton-RDNA4 \
  --revision v1.0.0 --include 'overlay/*' \
  --local-dir ./paiton-ornith-hf

Use the quick-start Docker command above, adding these two options before the image name:

  -e PAITON_ORNITH_OVERLAY=/models/paiton-overlay \
  -v "$PWD/paiton-ornith-hf/overlay:/models/paiton-overlay:ro" \

The runtime validates the artifact hashes before use. Keep the overlay/ directory intact: it contains the exact files expected by the release loader. The .so is loaded by Paiton; it is not a standalone Transformers checkpoint.

For an immutable runtime reference, replace the friendly image tag with:

ghcr.io/eliovp/paiton-vllm-plugin@sha256:f8feb0ea85e36f681eaf4f6c1d534551e4e0d1d98bf479f1cb6b6236a5893de2

File hashes and provenance are in SHA256SUMS and paiton-hub.json. All artifacts are byte-for-byte copies of the published release. This repository contains no checkpoint weights, draft weights, calibration data or compiler sources.

Links and licenses

The Paiton release is distributed under Apache-2.0, with retained third-party notices and license texts. Original model and draft terms remain those of their respective upstream repositories. This artifact package does not change those terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for EliovpAI/Ornith-1.5-35B-A3B-MXFP4-Paiton-RDNA4