--- license: apache-2.0 base_model: Qwen/Qwen2.5-32B-Instruct tags: - cascadia - openvino - int4 - pipeline-parallel library_name: cascadia --- # Qwen2.5 32B Instruct — Cascadia int4 shards Pre-exported [Cascadia](https://github.com/labscommunity/cascadia) inference shards for `Qwen/Qwen2.5-32B-Instruct`, so operators can deploy without running an export themselves. These are **OpenVINO IR** artifacts for Cascadia's `ov-runtime` engine. They are not loadable by `transformers` — see *Using these* below. ## Presets | Preset | Path | Stages | Layers | Size | |---|---|---|---|---| | `1` | `int4/stages-1` | 1 | 64 | 16.9 GB | | `2` | `int4/stages-2` | 2 | 64 | 16.9 GB | | `4` | `int4/stages-4` | 4 | 64 | 16.9 GB | Each preset is a separate export: a dense model's layer split is fixed at export time, so a 2-stage tree cannot serve a 4-node pipeline. Pick the one matching your fleet. ``` int4/stages-N/ pipeline_config.json # model geometry: layers, heads, rope, arch tag stage_0/ ... stage_N-1/ # each: openvino_model.xml/.bin + stage_config.json tokenizer/ # tokenizer.json + configs ``` `cascadia.json` at the repo root is the machine-readable index (sizes, checksums, export version) that Cascadia's model registry reads. ## Using these ```bash hf download communitylabs/cascadia-qwen2.5-32b-int4 --local-dir ./cascadia-qwen2.5-32b-int4 cascadia worker --model ./cascadia-qwen2.5-32b-int4/int4/stages-1 --engine ov-runtime ``` The worker takes a **local path**, never a HuggingFace id — Cascadia workers never download or convert models at serve time. ## Provenance Exported with Cascadia's `tools/export_shards.py`, `export_version: v5_canonical_inputs`, int4 via NNCF (group size 128). Quantization is lossy. For anything quality-sensitive, benchmark against the source model rather than assuming parity. ## License Apache-2.0, inherited from [`Qwen/Qwen2.5-32B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-32B-Instruct). These shards are a derivative work; the upstream license and its terms apply.