--- license: apache-2.0 library_name: coremltools pipeline_tag: zero-shot-classification base_model: fastino/GLiNER2.5-Decide language: - en tags: - coreml - gliner - gliner2 - deberta-v3 - apple-silicon - fp16 - multifunction - zero-shot-classification --- # GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 [fastino/GLiNER2.5-Decide](https://huggingface.co/fastino/GLiNER2.5-Decide) (340M, DeBERTa-v3-large, zero-shot classification) packaged for fast local inference on Apple Silicon: - **Core ML, FP16**: the classification path (encoder + label classifier), converted with the tooling from [FluidInference/gliner2-5-decide-coreml](https://huggingface.co/FluidInference/gliner2-5-decide-coreml). - **MultiFn**: one multifunction package with four fixed-length graphs (`L64`, `L128`, `L256`, `L512`) that share a single copy of the weights: **0.9 GB instead of 3.5 GB** for four separate packages. - **Routing**: every request is tokenized and sent to the smallest bucket that fits it, so short requests pay short-request latency. Text longer than 512 tokens is split into chunks automatically. ## Results (Apple M1 Pro, macOS 27.2, `CPU_AND_GPU`) | Bucket | Model only (Swift) | End to end (Python router) | `ALL` compute units | |---:|---:|---:|---:| | 64 | 19.3 ms | 20.0 ms | 70.7 ms | | 128 | 31.6 ms | 32.2 ms | 197.6 ms | | 256 | 63.3 ms | 60.9 ms | 79.6 ms | | 512 | 133.9 ms | 135.3 ms | 208.9 ms | p50 latency. "End to end" includes tokenization, routing, prediction and decoding; Python preprocessing costs only 0.2–0.7 ms, so a native Swift runtime would not be meaningfully faster. **Use the GPU**: on this machine `ALL` (GPU + Neural Engine) is slower at every size, and the Neural Engine alone is ~10x slower (see [Research notes](#research-notes)). **Fidelity.** Every bucket was verified against the native PyTorch model at conversion time (logit error ≤ 4.8e-7 for the export wrapper; identical labels on every example that fits). The test suite checks the router against native answers on six requests covering all four buckets (1–4 heads, multi-label, label descriptions, prompts): identical labels, largest confidence difference **0.0007**. ## Quick start Requires Apple Silicon, **macOS 15+** (multifunction models), Python 3.10–3.12 and [uv](https://docs.astral.sh/uv/). ```bash hf download augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 --local-dir GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 cd GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 uv sync ``` ```python from gliner_decide_coreml import DecideRouter router = DecideRouter() # loads all four buckets once; do this at startup router.classify( "My subscription renewed after the service was already down. Can I get that charge refunded?", { "intent": ["order_status", "refund_request", "cancel_subscription", "other"], "urgency": ["low", "normal", "high"], "topics": {"labels": ["billing", "outage", "account"], "multi_label": True, "cls_threshold": 0.4}, }, ) # {"intent": {"label": "refund_request", "confidence": 0.99}, "urgency": {...}, "topics": [{...}, ...]} result, route = router.classify_with_route(text, tasks) # Route(tokens=63, bucket=64, chunks=1, calls=1) ``` The task format is the same as native `classify_text`: a list of labels, or a dict with `labels` (list, or `{label: description}`), `multi_label`, `cls_threshold` and `prompt`. The first load of each function compiles it for the GPU (can take a minute or more on a new machine); macOS caches the result, and later loads take about 10 s for all four. ## How it works ``` text + tasks ──► GLiNER2 schema tokenizer ──► n tokens ──► smallest bucket ≥ n ──► Core ML function ──► decode "( [P] intent ( [L] refund …" 64 / 128 / 256 / 512 (GPU, fp16) softmax or [SEP_TEXT] my subscription …" └─ n > 512: word chunks, sigmoid + merged per head threshold ``` - **Input budget.** The schema (questions and labels) shares the sequence with the text. With 2 questions and 8 labels (38 tokens) the buckets leave room for about 26 / 90 / 218 / 474 text tokens. Label descriptions and prompts cost tokens too. - **Heads.** Each call scores up to 4 heads with up to 32 labels each. More than 4 heads are split across calls. Heads in one call see each other's labels, so results can shift slightly from a single native pass. - **Chunking.** For text over 512 tokens: overlapping word chunks (32 words of overlap), each classified, then merged: single-label heads take the most confident chunk, multi-label heads take the per-label maximum. This is a heuristic; no chunk sees the whole text. ## Layout ``` models/ GLiNER2.5-Decide-MultiFn-fp16.mlpackage functions L64 / L128 / L256 / L512 (default L128) tokenizer/ DeBERTa-v3 tokenizer + GLiNER2 special tokens src/gliner_decide_coreml/ router.py DecideRouter: bucket selection, head grouping, chunking encoding.py GLiNER2 schema preprocessing and bucket padding decoding.py activations, thresholds, chunk merging tests/ router vs native answers (golden.json), chunking, head splitting scripts/ make_golden.py record native PyTorch answers for the tests build/convert_bucket.py convert + verify one bucket (from the Fluid Inference tooling) build/build_multifunction.py bench/ Swift latency and compute-plan tools, Python router benchmark research/ane-ceiling/ Neural Engine feasibility experiment (see below) ``` ## Tests and benchmarks ```bash uv run pytest # router vs native answers, all buckets uv run python bench/router_latency.py # end-to-end latency per bucket cd bench swiftc -O benchbuckets.swift -o benchbuckets # model-only latency of compiled packages swiftc -O plan.swift -o plan # per-op device placement (MLComputePlan) ``` `scripts/make_golden.py` regenerates `tests/golden.json` from the native model (downloads the pinned checkpoint, ~1.7 GB). ## Rebuilding the package ```bash cd scripts/build for L in 64 128 256 512; do uv run python convert_bucket.py --length $L --max-heads 4 --max-options 32 --output-dir build done uv run python build_multifunction.py --build-dir build ``` Source checkpoint: `fastino/GLiNER2.5-Decide` at revision `65624f1a0265b3f612bae66a2685a06b94a68a9d`. Pinned toolchain: coremltools 9.0, torch 2.7.0, transformers 4.57.6, gliner2 2.0.0. ## Research notes Measured on an M1 Pro while building this package. - **Compute units.** Under `ALL`, Core ML splits every layer between the GPU and the Neural Engine (48 switches per request), which is slower than the GPU alone. Under `CPU_AND_NEURAL_ENGINE` the model takes ~600 ms: DeBERTa's disentangled attention uses `gather_along_axis` (49 ops: the c2p and p2c relative-position lookups in all 24 layers, plus the label-marker gather), which the Neural Engine cannot run. - **Neural Engine ceiling** (`research/ane-ceiling`, random-weight encoders with this model's shape). A plain encoder runs in **32–36 ms** on the Neural Engine, and adding DeBERTa's extra position matmuls costs only +6 ms. Replacing the gathers exactly with gather-free "relative shift" code keeps every op on the Neural Engine but is slow: reshape-based shift **807 ms**, barrel shifter (static slices + blends) **187 ms**. An exact ANE rewrite would only beat the GPU if the shift cost less than ~1 ms per layer. - **GPU headroom.** The model needs roughly 175 GFLOP per 256-token request; at the M1 Pro GPU's peak that is a ~35 ms floor, versus 61–63 ms achieved. Bigger wins come from shorter inputs (this package) or a smaller, distilled model, not from kernel tuning. ## Limitations - Classification only: GLiNER2's entity, relation and structure extraction heads are not included. - Preprocessing uses the GLiNER2 Python package, which imports PyTorch (no model weights are loaded). - Latency figures are from one machine (M1 Pro); newer chips are faster and may favour different compute units. - Multifunction packages need macOS 15 / iOS 18 or later. ## License and attribution Apache-2.0 (see `LICENSE` and `NOTICE`). Model by [Fastino](https://huggingface.co/fastino/GLiNER2.5-Decide); original Core ML conversion tooling by [Fluid Inference](https://huggingface.co/FluidInference/gliner2-5-decide-coreml).