NEO CODER MAX 27B: native GGUF through vLLM on AMD RDNA4

DavidAU's NEO CODER MAX fine-tune, served through vLLM and accelerated by native Paiton kernels. Code, chat and ask questions about images through an OpenAI-compatible API. The original mixed Q4_K_M GGUF weights retain their quantized values.

Complete streaming requests had 6.4%, 5.1% and 0.8% lower median latency than llama.cpp in the three matched 128-output text workloads below. Results apply to this model and profile; some prefill-only cases favor llama.cpp.

This repository packages Paiton v1.1.0 runtime artifacts and metadata, about 28.4 MB in overlay/. The fine-tune and GGUF are by DavidAU. Paiton adds compiled execution; no additional training or weight quantization is performed here. Weights and the image projector download directly from the pinned upstream repository. The compiler stays private and is not needed to run the package.

Run locally

Use Linux x86-64, Docker and a Radeon AI PRO R9700 with working AMD GPU device access. The pinned image includes the qualified ROCm 7.14 and vLLM runtime.

docker run -d --name paiton-qwen38-neo \
  --device /dev/kfd --device /dev/dri --group-add video --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v paiton-qwen38-neo-cache:/models/cache \
  ghcr.io/eliovp/paiton-vllm-plugin@sha256:534287969135f581744ae481b578599468b0bf7ac9a4051b0941500e4c18da4d

The first start downloads approximately 19.43 GB of language weights and the projector. Later starts reuse the cache and verify the file hashes. Follow docker logs -f paiton-qwen38-neo until startup completes, then check curl --fail http://127.0.0.1:8000/health.

curl --fail http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"qwen38-neo","messages":[{"role":"user","content":"Write a Python function that removes duplicate integers while preserving their order."}],"temperature":0,"max_tokens":256,"stream":true,"chat_template_kwargs":{"enable_thinking":false}}'

Set chat_template_kwargs.enable_thinking explicitly to choose reasoning behavior. Image requests use the same endpoint with text and image_url content items; PNG/JPEG URLs and base64 data URLs are supported. Image request example.

Measured performance

Same R9700, pinned GGUF, tokenizer, greedy sampling, 8K context, BF16 KV, one active sequence and fixed token counts. Each workload uses one warmup and five measured requests. Times include the complete streaming HTTP request; p95 is interpolated from five samples.

Input / output tokens llama.cpp median / p95 (s) Paiton + vLLM median / p95 (s) Median latency reduction
128 / 128 5.264 / 5.289 4.925 / 4.929 6.4%
1,024 / 128 5.933 / 5.936 5.631 / 5.637 5.1%
4,096 / 128 9.099 / 9.101 9.023 / 9.030 0.8%

This is a comparison with working llama.cpp, not stock vLLM. It does not establish a speed advantage for every prompt or deployment. Short and long prefill-only requests favor llama.cpp. The engines use different activation arithmetic; bit-identical cross-engine output is not claimed. Full text/image results, quality checks and reproduction.

The fixed task suite scored 10/11 in both engines, with the same code-trace failure; five image fixtures and JPEG checks passed. The held-out Q8 decode evaluation measured a 0.1301% perplexity increase against its FP32 decode control over 770 predictions, within the preset 1% bound. These are bounded checks, not a broad accuracy claim. This HF package uses the published artifacts unchanged and introduces no new performance measurements.

Supported profile and precision

Setting Qualified contract
Hardware Radeon AI PRO R9700, gfx1201, 32 GB
Context 8,192 total tokens, including image and output tokens
Requests One active sequence; additional HTTP requests queue; TP1
Input Text and one PNG/JPEG image; up to 1,024 image embeddings
API Chat, completions, streaming, Qwen3 reasoning and Qwen3 Coder tool parsing
MTP / prefix caching / video Disabled
Execution Native HIP language and vision artifacts; native decode graphs; vLLM loader, scheduler and sampling

The language file contains Q4_K and Q6_K matrices, FP32 coefficients, a BF16 output head and unused Q8_0 MTP tensors. Paiton preserves the packed weights. Large Q4 decode projections use Q8_1 activation operands; prefill matrix operands use FP16 with FP32 accumulation. GDN state and image embeddings remain FP32, and KV storage is BF16. The tokenizer, chat template, special tokens and stop behavior come from the pinned author source.

Sampled peak GPU use was 23.736 GiB for Paiton and 19.034 GiB for llama.cpp across the reported qualification workloads. These figures include runtime allocations and are not minimum-VRAM guarantees. Smaller GPUs are not qualified. Complete model and arithmetic contract.

Download the HF artifacts

The Docker image already contains these artifacts. To use this repository's exact payload instead, download it with the Hugging Face CLI:

hf download EliovpAI/Qwen3.8-27B-NEO-CODER-MAX-Q4_K_M-GGUF-Paiton-RDNA4 \
  --revision v1.1.0 --include 'overlay/*' \
  --local-dir ./paiton-neo-hf

Then start the image with the downloaded directory mounted read-only. Run only one of the two launch examples at a time; docker stop paiton-qwen38-neo stops the first example.

docker run --rm --name paiton-qwen38-neo-hf \
  --device /dev/kfd --device /dev/dri --group-add video --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v paiton-qwen38-neo-cache:/models/cache \
  --mount "type=bind,src=$PWD/paiton-neo-hf/overlay,dst=/models/paiton-overlay,readonly" \
  ghcr.io/eliovp/paiton-vllm-plugin@sha256:534287969135f581744ae481b578599468b0bf7ac9a4051b0941500e4c18da4d \
  --model-dir /models/paiton-overlay

Keep overlay/ intact: startup validates its inventory and hashes. These shared libraries load through Paiton's runtime; this is not a standalone Transformers weight checkpoint. To return to the image's bundled artifacts, stop this container and use the first launch command. Both commands reuse the same weight cache.

Provenance and licenses

Paiton runtime artifacts use Apache-2.0, with retained third-party notices and license texts. The upstream model repositories declare Apache-2.0; their attribution and terms remain applicable. llama.cpp is a benchmark reference, not a serving dependency. The compiled Paiton model uses native HIP/C++; the external vLLM integration retains its existing framework dependencies.

Explore Paiton · Optimized-model library

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support