Instructions to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l") model = AutoModelForMultimodalLM.from_pretrained("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l
- SGLang
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l with Docker Model Runner:
docker model run hf.co/littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l
Prismyra decision FP8, 36 layers, for Qwen3.6-35B-A3B-FP8
A full-weight FP8 checkpoint truncated to 36 of the base model's 40 layers, with a LoRA trained specifically for this depth -- distilled from the 40-layer checkpoint's own output distribution -- folded into the FP8 weights and requantized per 128x128 block. It is for Prismyra: an engine that answers typed questions (booleans, choices) about a document by reading the probability of each declared option's token in a single forward pass, rather than generating an answer -- see the Prismyra repository for what that trades off and where it does not apply.
Of the three FP8 checkpoints in this family (40, 36, 32 layers), this is the one recommended by default: on the five sets below it ties or clears every bar the 40-layer checkpoint does, while running measurably faster.
This revision raises the folded-in LoRA's rank from 16 to 64 (alpha scaled from 32 to 128 to keep the same
effective update magnitude) and retrains it by knowledge distillation from this checkpoint's own 36-layer forward
pass. The previous revision (rank-16 LoRA) stays available at commit
e24a3c89e7185f28f02f8ebca878804bbffe768d;
every absolute number quoted below for "the previous revision" is that commit's own published numbers, not re-measured here.
Use
pip install "prismyra[server,fast] @ git+https://github.com/littlemex/Prismyra@v0.3.0"
from prismyra import Prismyra, Boolean
engine = Prismyra("littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l")
result = engine.ask(document, [Boolean(id="q", prompt="Is this a two-sided agreement?")])
prismyra-serve --model littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l --require-kernels
--require-kernels refuses to start rather than serve at a fraction of the speed. Prismyra's fused kernels are
matched against the checkpoint by per-layer tensor shape, not by layer count or adapter rank, so this checkpoint gets
every kernel the base model gets regardless of which revision's LoRA is folded in -- there is no flag to set for this.
Training
The adapter targets 250 modules (attention and gated-delta-net projections, the shared expert) at rank 64 instead of the previous revision's 16 (76.8M parameters instead of 19.2M), initialised by zero-padding the previous revision's own trained rank-16 adapter up to rank 64 so training continues from where that adapter already was rather than from zero. Training data is the same pool used for the previous revision's own adapter (RACE, SciQ, BoolQ-family and twelve others, 31,418 rows); the distillation target is this checkpoint's own 36-layer forward pass, with the gold label on about 7% of rows -- the rows where the pool's own labeled answer and a frontier LLM's (Kimi K3) answer disagreed -- taken from that frontier LLM's answer letter rather than the pool's own label. One epoch, learning rate 2.5e-5, three L40S GPUs, about 19 GPU-hours.
On 2,500 held-out training-distribution rows, this reduced the KL divergence to the distillation target by 51.2%, against 26.8% for the previous revision's own rank-16 adapter measured the same way -- the only point in this project's rank progression (16, 64, 128) where raising the rank measurably helped on this axis; rank 128, tried next, did not improve on this result on any of three held-out checks and was not published.
Evaluation
Five sets, served by Prismyra on one L40S, compared against the previous revision (rank-16 LoRA, commit
e24a3c89), the same protocol as that revision's own published numbers: RACE (579 questions), BoolQ (400), bury7k
and bury10k (the same RACE questions buried in about 7,000 and 10,000 tokens of unrelated articles, 157 questions
each), and Kev (a 764-question transfer test, of which about 30% comes from families this checkpoint never trained on
in any form).
| set | point difference vs. the previous revision | 95% interval | sign-test p |
|---|---|---|---|
| RACE (579) | +0.17 | [-0.69, +1.04] | -- |
| BoolQ (400) | +0.25 | [-0.75, +1.25] | -- |
| bury7k (157) | 0.00 | [-3.18, +3.18] | -- |
| bury10k (157) | +1.27 | [-1.27, +4.46] | -- |
| Kev (764) | +0.92 | [-0.39, +2.36] | 0.265 |
This is the first point in this project's rank progression with a point estimate at or above the previous revision on every one of these five sets. It is also the first to clear this project's weaker, 2,050-question held-out set (disjoint from every set above and from this checkpoint's own training data) by a margin a sign test calls significant: +1.51 points, 95% interval [+0.49, +2.59], p = 0.0062. Read together with the row above: the margin is positive and significant on a broad held-out slice, but on this project's specific production bar -- Kev, the 764-question set, sign-test p < 0.05 -- it is positive but not significant (p = 0.265).
bury7k is the plainest result to read on its own: zero point difference on 157 questions buried in roughly 7,000 tokens of unrelated text, so the higher-rank adapter shows no advantage on that specific long-context slice even where bury10k, 3,000 tokens longer, shows one.
This checkpoint is also measured against JevBench (a 231-question benchmark none of these checkpoints trained on) and against four general-purpose LLMs on the same five sets above, in the training recipe's README. The short version, measured on the previous (rank-16) revision and not re-measured for this one: this checkpoint's family is not significantly different from the 40-layer checkpoint on JevBench, but its calibration there (Brier and expected calibration error) is worse than the 40-layer checkpoint's; and against general-purpose LLMs it loses on accuracy overall, most clearly on Kev, at one to a few orders of magnitude lower estimated GPU-time cost depending on how long the context is.
Latency
Two conditions. First, this revision against the previous revision directly, same L40S, same script, measured twice in alternation to separate a real difference from GPU/driver warm-up state:
| 1 question | 16 questions | 64 questions | |
|---|---|---|---|
| previous revision (rank-16 LoRA) | 233.7 ms | 348.4 ms | 623.1 ms |
| this revision (rank-64 LoRA) | 233.3 ms | 348.4 ms | 622.8 ms |
Within 0.2 ms at every width -- folding in a higher-rank adapter costs nothing measurable, which is expected: the adapter changes weight values, not the checkpoint's shape or layer count, so it gets the same kernels at the same speed.
Second, measured on the previous revision against the 40- and 32-layer checkpoints and the untrained base, same
script and GPU, on the same ~5,300-token document repeated across calls (repeating a document, once warmed up, costs
the same as making it fresh on every call -- see docs/PERFORMANCE.md in the repository). Not re-measured for this
revision, but the comparison above shows adapter rank does not move these numbers:
| checkpoint | 1 question | 16 questions | 64 questions |
|---|---|---|---|
| base, untrained, 40 layers | 276.8 ms | 408.7 ms | 733.4 ms |
| 40-layer | 269.9 ms | 399.4 ms | 716.8 ms |
| 36-layer (this family) | 244.3 ms | 362.1 ms | 650.3 ms |
| 32-layer | 218.0 ms | 324.5 ms | 584.7 ms |
The speed difference across the three published checkpoints tracks their layer count, which is what fused kernels applying at every depth looks like -- a checkpoint that fell back to the unfused path would be roughly three times slower than the figures above, at any depth.
How it relates to the other checkpoints
This project publishes five repositories:
littlemex/prismyra-decision-lora-qwen3.6-35b-a3b-- the LoRA adapter for the full 40-layer base model, trained separately from this checkpoint's own LoRA.littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-40l-- all 40 layers, that adapter folded in. The teacher this checkpoint's adapter was distilled from.littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-36l-- this repository.littlemex/prismyra-decision-qwen3.6-35b-a3b-fp8-32l-- 32 layers, rank-16 adapter. Faster again, and ahead of both longer checkpoints on the two long-context sets above, but the one of the three that misses the BoolQ bar (by a single question) and is not significantly different from the untrained base on JevBench.littlemex/prismyra-decision-qwen3.6-35b-a3b-nvfp4-36l-- this checkpoint's routed experts in NVFP4 instead of FP8, for Blackwell GPUs.
Licence and data terms
Apache-2.0, the base model's licence, and this repository carries the base model's LICENSE file unchanged. Some of
the training sources behind this checkpoint's LoRA carry their own terms -- RACE and SciQ are distributed for
non-commercial research use -- so check those terms before commercial use of this checkpoint. About 7% of training
rows (this revision only; the previous revision's adapter did not use this) take their gold label from a frontier LLM
(Kimi K3, called through its provider's API) rather than from the pool's own label, used only as a single-letter
answer where it disagreed with the pool; no generated rationale or free text from that model is reproduced in this
checkpoint's weights or in this card.
- Downloads last month
- 325