How to use from the
Use from the
GLiNER2 library
from gliner2 import GLiNER2

model = GLiNER2.from_pretrained("augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512")

# Extract entities
text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday."
result = extractor.extract_entities(text, ["company", "person", "product", "location"])

print(result)

GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512

fastino/GLiNER2.5-Decide (340M, DeBERTa-v3-large, zero-shot classification) packaged for fast local inference on Apple Silicon:

  • Core ML, FP16: the classification path (encoder + label classifier), converted with the tooling from FluidInference/gliner2-5-decide-coreml.
  • MultiFn: one multifunction package with four fixed-length graphs (L64, L128, L256, L512) that share a single copy of the weights: 0.9 GB instead of 3.5 GB for four separate packages.
  • Routing: every request is tokenized and sent to the smallest bucket that fits it, so short requests pay short-request latency. Text longer than 512 tokens is split into chunks automatically.

Results (Apple M1 Pro, macOS 27.2, CPU_AND_GPU)

Bucket Model only (Swift) End to end (Python router) ALL compute units
64 19.3 ms 20.0 ms 70.7 ms
128 31.6 ms 32.2 ms 197.6 ms
256 63.3 ms 60.9 ms 79.6 ms
512 133.9 ms 135.3 ms 208.9 ms

p50 latency. "End to end" includes tokenization, routing, prediction and decoding; Python preprocessing costs only 0.2–0.7 ms, so a native Swift runtime would not be meaningfully faster. Use the GPU: on this machine ALL (GPU + Neural Engine) is slower at every size, and the Neural Engine alone is ~10x slower (see Research notes).

Fidelity. Every bucket was verified against the native PyTorch model at conversion time (logit error ≤ 4.8e-7 for the export wrapper; identical labels on every example that fits). The test suite checks the router against native answers on six requests covering all four buckets (1–4 heads, multi-label, label descriptions, prompts): identical labels, largest confidence difference 0.0007.

Quick start

Requires Apple Silicon, macOS 15+ (multifunction models), Python 3.10–3.12 and uv.

hf download augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512 --local-dir GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512
cd GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512
uv sync
from gliner_decide_coreml import DecideRouter

router = DecideRouter()  # loads all four buckets once; do this at startup

router.classify(
    "My subscription renewed after the service was already down. Can I get that charge refunded?",
    {
        "intent": ["order_status", "refund_request", "cancel_subscription", "other"],
        "urgency": ["low", "normal", "high"],
        "topics": {"labels": ["billing", "outage", "account"], "multi_label": True, "cls_threshold": 0.4},
    },
)
# {"intent": {"label": "refund_request", "confidence": 0.99}, "urgency": {...}, "topics": [{...}, ...]}

result, route = router.classify_with_route(text, tasks)
# Route(tokens=63, bucket=64, chunks=1, calls=1)

The task format is the same as native classify_text: a list of labels, or a dict with labels (list, or {label: description}), multi_label, cls_threshold and prompt.

The first load of each function compiles it for the GPU (can take a minute or more on a new machine); macOS caches the result, and later loads take about 10 s for all four.

How it works

text + tasks ──► GLiNER2 schema tokenizer ──► n tokens ──► smallest bucket ≥ n ──► Core ML function ──► decode
                 "( [P] intent ( [L] refund …"             64 / 128 / 256 / 512     (GPU, fp16)          softmax or
                  [SEP_TEXT] my subscription …"             └─ n > 512: word chunks,                    sigmoid +
                                                               merged per head                          threshold
  • Input budget. The schema (questions and labels) shares the sequence with the text. With 2 questions and 8 labels (38 tokens) the buckets leave room for about 26 / 90 / 218 / 474 text tokens. Label descriptions and prompts cost tokens too.
  • Heads. Each call scores up to 4 heads with up to 32 labels each. More than 4 heads are split across calls. Heads in one call see each other's labels, so results can shift slightly from a single native pass.
  • Chunking. For text over 512 tokens: overlapping word chunks (32 words of overlap), each classified, then merged: single-label heads take the most confident chunk, multi-label heads take the per-label maximum. This is a heuristic; no chunk sees the whole text.

Layout

models/
  GLiNER2.5-Decide-MultiFn-fp16.mlpackage   functions L64 / L128 / L256 / L512 (default L128)
  tokenizer/                                DeBERTa-v3 tokenizer + GLiNER2 special tokens
src/gliner_decide_coreml/
  router.py      DecideRouter: bucket selection, head grouping, chunking
  encoding.py    GLiNER2 schema preprocessing and bucket padding
  decoding.py    activations, thresholds, chunk merging
tests/           router vs native answers (golden.json), chunking, head splitting
scripts/
  make_golden.py             record native PyTorch answers for the tests
  build/convert_bucket.py    convert + verify one bucket (from the Fluid Inference tooling)
  build/build_multifunction.py
bench/           Swift latency and compute-plan tools, Python router benchmark
research/ane-ceiling/        Neural Engine feasibility experiment (see below)

Tests and benchmarks

uv run pytest                          # router vs native answers, all buckets
uv run python bench/router_latency.py  # end-to-end latency per bucket

cd bench
swiftc -O benchbuckets.swift -o benchbuckets   # model-only latency of compiled packages
swiftc -O plan.swift -o plan                   # per-op device placement (MLComputePlan)

scripts/make_golden.py regenerates tests/golden.json from the native model (downloads the pinned checkpoint, ~1.7 GB).

Rebuilding the package

cd scripts/build
for L in 64 128 256 512; do
  uv run python convert_bucket.py --length $L --max-heads 4 --max-options 32 --output-dir build
done
uv run python build_multifunction.py --build-dir build

Source checkpoint: fastino/GLiNER2.5-Decide at revision 65624f1a0265b3f612bae66a2685a06b94a68a9d. Pinned toolchain: coremltools 9.0, torch 2.7.0, transformers 4.57.6, gliner2 2.0.0.

Research notes

Measured on an M1 Pro while building this package.

  • Compute units. Under ALL, Core ML splits every layer between the GPU and the Neural Engine (48 switches per request), which is slower than the GPU alone. Under CPU_AND_NEURAL_ENGINE the model takes ~600 ms: DeBERTa's disentangled attention uses gather_along_axis (49 ops: the c2p and p2c relative-position lookups in all 24 layers, plus the label-marker gather), which the Neural Engine cannot run.
  • Neural Engine ceiling (research/ane-ceiling, random-weight encoders with this model's shape). A plain encoder runs in 32–36 ms on the Neural Engine, and adding DeBERTa's extra position matmuls costs only +6 ms. Replacing the gathers exactly with gather-free "relative shift" code keeps every op on the Neural Engine but is slow: reshape-based shift 807 ms, barrel shifter (static slices + blends) 187 ms. An exact ANE rewrite would only beat the GPU if the shift cost less than ~1 ms per layer.
  • GPU headroom. The model needs roughly 175 GFLOP per 256-token request; at the M1 Pro GPU's peak that is a ~35 ms floor, versus 61–63 ms achieved. Bigger wins come from shorter inputs (this package) or a smaller, distilled model, not from kernel tuning.

Limitations

  • Classification only: GLiNER2's entity, relation and structure extraction heads are not included.
  • Preprocessing uses the GLiNER2 Python package, which imports PyTorch (no model weights are loaded).
  • Latency figures are from one machine (M1 Pro); newer chips are faster and may favour different compute units.
  • Multifunction packages need macOS 15 / iOS 18 or later.

License and attribution

Apache-2.0 (see LICENSE and NOTICE). Model by Fastino; original Core ML conversion tooling by Fluid Inference.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for augustoFranke/GLiNER2.5-Decide-CoreML-FP16-MultiFn-L64-512

Quantized
(8)
this model