Instructions to use augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER
How to use augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512") - GLiNER2
How to use augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512
fastino/GLiNER2.5-Decide (340M, DeBERTa-v3-large, zero-shot classification) packaged for fast local inference on Apple Silicon:
- Core ML, FP16: the classification path (encoder + label classifier), converted with the tooling from FluidInference/gliner2-5-decide-coreml.
- MultiFn: one multifunction package with four fixed-length graphs (
L64,L128,L256,L512) that share a single copy of the weights: 0.9 GB instead of 3.5 GB for four separate packages. - Routing: every request is tokenized and sent to the smallest bucket that fits it, so short requests pay short-request latency. Text longer than 512 tokens is split into chunks automatically.
Results (Apple M1 Pro, macOS 27.2, CPU_AND_GPU)
| Bucket | Model only (Swift) | End to end (Python router) | ALL compute units |
|---|---|---|---|
| 64 | 19.3 ms | 20.0 ms | 70.7 ms |
| 128 | 31.6 ms | 32.2 ms | 197.6 ms |
| 256 | 63.3 ms | 60.9 ms | 79.6 ms |
| 512 | 133.9 ms | 135.3 ms | 208.9 ms |
p50 latency. "End to end" includes tokenization, routing, prediction and decoding; Python preprocessing costs
only 0.2–0.7 ms, so a native Swift runtime would not be meaningfully faster. Use the GPU: on this machine
ALL (GPU + Neural Engine) is slower at every size, and the Neural Engine alone is ~10x slower (see
Research notes).
Fidelity. Every bucket was verified against the native PyTorch model at conversion time (logit error ≤ 4.8e-7 for the export wrapper; identical labels on every example that fits). The test suite checks the router against native answers on six requests covering all four buckets (1–4 heads, multi-label, label descriptions, prompts): identical labels, largest confidence difference 0.0007.
Quick start
Requires Apple Silicon, macOS 15+ (multifunction models), Python 3.10–3.12 and uv.
hf download augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 --local-dir GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512
cd GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512
uv sync
from gliner_decide_coreml import DecideRouter
router = DecideRouter() # loads all four buckets once; do this at startup
router.classify(
"My subscription renewed after the service was already down. Can I get that charge refunded?",
{
"intent": ["order_status", "refund_request", "cancel_subscription", "other"],
"urgency": ["low", "normal", "high"],
"topics": {"labels": ["billing", "outage", "account"], "multi_label": True, "cls_threshold": 0.4},
},
)
# {"intent": {"label": "refund_request", "confidence": 0.99}, "urgency": {...}, "topics": [{...}, ...]}
result, route = router.classify_with_route(text, tasks)
# Route(tokens=63, bucket=64, chunks=1, calls=1)
The task format is the same as native classify_text: a list of labels, or a dict with labels (list, or
{label: description}), multi_label, cls_threshold and prompt.
The first load of each function compiles it for the GPU (can take a minute or more on a new machine); macOS caches the result, and later loads take about 10 s for all four.
How it works
text + tasks ──► GLiNER2 schema tokenizer ──► n tokens ──► smallest bucket ≥ n ──► Core ML function ──► decode
"( [P] intent ( [L] refund …" 64 / 128 / 256 / 512 (GPU, fp16) softmax or
[SEP_TEXT] my subscription …" └─ n > 512: word chunks, sigmoid +
merged per head threshold
- Input budget. The schema (questions and labels) shares the sequence with the text. With 2 questions and 8 labels (38 tokens) the buckets leave room for about 26 / 90 / 218 / 474 text tokens. Label descriptions and prompts cost tokens too.
- Heads. Each call scores up to 4 heads with up to 32 labels each. More than 4 heads are split across calls. Heads in one call see each other's labels, so results can shift slightly from a single native pass.
- Chunking. For text over 512 tokens: overlapping word chunks (32 words of overlap), each classified, then merged: single-label heads take the most confident chunk, multi-label heads take the per-label maximum. This is a heuristic; no chunk sees the whole text.
Layout
models/
GLiNER2.5-Decide-MultiFn-fp16.mlpackage functions L64 / L128 / L256 / L512 (default L128)
tokenizer/ DeBERTa-v3 tokenizer + GLiNER2 special tokens
src/gliner_decide_coreml/
router.py DecideRouter: bucket selection, head grouping, chunking
encoding.py GLiNER2 schema preprocessing and bucket padding
decoding.py activations, thresholds, chunk merging
tests/ router vs native answers (golden.json), chunking, head splitting
scripts/
make_golden.py record native PyTorch answers for the tests
build/convert_bucket.py convert + verify one bucket (from the Fluid Inference tooling)
build/build_multifunction.py
bench/ Swift latency and compute-plan tools, Python router benchmark
research/ane-ceiling/ Neural Engine feasibility experiment (see below)
Tests and benchmarks
uv run pytest # router vs native answers, all buckets
uv run python bench/router_latency.py # end-to-end latency per bucket
cd bench
swiftc -O benchbuckets.swift -o benchbuckets # model-only latency of compiled packages
swiftc -O plan.swift -o plan # per-op device placement (MLComputePlan)
scripts/make_golden.py regenerates tests/golden.json from the native model (downloads the pinned
checkpoint, ~1.7 GB).
Rebuilding the package
cd scripts/build
for L in 64 128 256 512; do
uv run python convert_bucket.py --length $L --max-heads 4 --max-options 32 --output-dir build
done
uv run python build_multifunction.py --build-dir build
Source checkpoint: fastino/GLiNER2.5-Decide at revision 65624f1a0265b3f612bae66a2685a06b94a68a9d.
Pinned toolchain: coremltools 9.0, torch 2.7.0, transformers 4.57.6, gliner2 2.0.0.
Research notes
Measured on an M1 Pro while building this package.
- Compute units. Under
ALL, Core ML splits every layer between the GPU and the Neural Engine (48 switches per request), which is slower than the GPU alone. UnderCPU_AND_NEURAL_ENGINEthe model takes ~600 ms: DeBERTa's disentangled attention usesgather_along_axis(49 ops: the c2p and p2c relative-position lookups in all 24 layers, plus the label-marker gather), which the Neural Engine cannot run. - Neural Engine ceiling (
research/ane-ceiling, random-weight encoders with this model's shape). A plain encoder runs in 32–36 ms on the Neural Engine, and adding DeBERTa's extra position matmuls costs only +6 ms. Replacing the gathers exactly with gather-free "relative shift" code keeps every op on the Neural Engine but is slow: reshape-based shift 807 ms, barrel shifter (static slices + blends) 187 ms. An exact ANE rewrite would only beat the GPU if the shift cost less than ~1 ms per layer. - GPU headroom. The model needs roughly 175 GFLOP per 256-token request; at the M1 Pro GPU's peak that is a ~35 ms floor, versus 61–63 ms achieved. Bigger wins come from shorter inputs (this package) or a smaller, distilled model, not from kernel tuning.
Limitations
- Classification only: GLiNER2's entity, relation and structure extraction heads are not included.
- Preprocessing uses the GLiNER2 Python package, which imports PyTorch (no model weights are loaded).
- Latency figures are from one machine (M1 Pro); newer chips are faster and may favour different compute units.
- Multifunction packages need macOS 15 / iOS 18 or later.
License and attribution
Apache-2.0 (see LICENSE and NOTICE). Model by Fastino;
original Core ML conversion tooling by Fluid Inference.
- Downloads last month
- 10