Instructions to use smdesai/GLiNER25-Decide-FP16-CoreML with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use smdesai/GLiNER25-Decide-FP16-CoreML with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("smdesai/GLiNER25-Decide-FP16-CoreML") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5-Decide β Core ML (iOS 18+ / macOS 15+)
| File | Purpose |
|---|---|
GLiNER25-Decide-FP16-CoreML.mlpackage/ |
Core ML package, FP16 weights and compute (~945 MB), three functions sharing one set of weights |
tokenizer.json, tokenizer_config.json, special_tokens_map.json |
Checkpoint tokenizer |
model-card.md |
Upstream model card (fastino/GLiNER2.5-Decide, Apache-2.0) |
Source: checkpoint revision bbe10ff77ebb238777c17d3a8ac9260e30929057,
classification-only (DeBERTa encoder + classifier). Built with coremltools
9.0 / Torch 2.7.1 from FP16 128/256/512 exports (finite FP16 attention
mask), merged into one multifunction package with shared weights.
Interface
Functions context128, context256, context512 (select with
MLModelConfiguration.functionName):
- inputs:
input_idsint32[1, L],attention_maskint32[1, L](1 = real token, 0 = padding) - output:
logits[1, L], one score per token position. Classification reads the positions of the schema's label markers ([L]).
Choose the smallest function that fits the request and pad to its length. The maximum is 512 tokens. Reject longer requests; do not truncate.
Required: GPU compute units
let configuration = MLModelConfiguration()
configuration.computeUnits = .cpuAndGPU
configuration.functionName = "context256"
let model = try MLModel(contentsOf: compiledURL, configuration: configuration)
Do not use INT8 weights on the GPU: they were nondeterministic and changed
decisions. Do not rely on .cpuOnly for this package either: FP16 on CPU
changes no decisions but exceeds the strict gate (max 0.012). Xcode compiles
the .mlpackage to .mlmodelc when it is bundled in an app; use
MLModel.compileModel(at:) otherwise.
Validation (iPhone 17 Pro, iOS 27.2, GPU)
Strict gate: every decision matches the FP32 PyTorch oracle and max probability error β€ 0.005. Two launches per function, identical results:
| Function | Corpus | Max error | Median |
|---|---|---|---|
| context128 | 75 | 0.001371 | 26 ms |
| context256 | 86 | 0.001371 | 44 ms |
| context512 | 95 (9 requests of 296β512 tokens) | 0.001371 | 121β124 ms |
Load β 1.7β2.0 s; physical footprint β€ 0.32 GB (weights are file-mapped). Not yet validated on an iOS 18β26 device or an older GPU.
SHA-256
bc1092e6eb7185c9ff1c3fb8e8659b70b8c12f295fc573c39c197f768c102078 GLiNER25-Decide-FP16-CoreML.mlpackage/Data/com.apple.CoreML/weights/weight.bin
fe39a4ea16cb1d35699d4cc2571f9c4e80692caccdafbb689645be56867f8d9f GLiNER25-Decide-FP16-CoreML.mlpackage/Data/com.apple.CoreML/model.mlmodel
3ad87d9ffe669147063e70850927dd2da90249e2acc5c8527f1eb65df467bcc8 tokenizer.json
- Downloads last month
- 9