Text Classification
Transformers.js
ONNX
GLiNER2
English
deberta-v2
feature-extraction
webgpu
typed-decisions
open-jev
system-one
Instructions to use onnx-community/GLiNER2.5-Decide-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use onnx-community/GLiNER2.5-Decide-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-classification', 'onnx-community/GLiNER2.5-Decide-ONNX'); - GLiNER2
How to use onnx-community/GLiNER2.5-Decide-ONNX with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("onnx-community/GLiNER2.5-Decide-ONNX") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
Download conversion/quantize_q4.py from onnx-community/GLiNER2.5-Decide-ONNX: direct link, hf CLI and curl.
- Browser
- Download file 1.64 kB
-
https://huggingface.co/onnx-community/GLiNER2.5-Decide-ONNX/resolve/main/conversion/quantize_q4.py
- Command line
-
hf download hf://onnx-community/GLiNER2.5-Decide-ONNX/conversion/quantize_q4.py
-
curl -L -o quantize_q4.py https://huggingface.co/onnx-community/GLiNER2.5-Decide-ONNX/resolve/main/conversion/quantize_q4.py
1.64 kB
| import onnx, os, time, numpy as np, sys | |
| from onnxconverter_common import float16 | |
| from onnxruntime.quantization import quantize_dynamic, QuantType | |
| from onnxruntime.quantization.matmul_nbits_quantizer import MatMulNBitsQuantizer | |
| D = "/Users/shreyas/work/rnd/gliner2/work/export/onnx" | |
| model = onnx.load(D + "/model.onnx") | |
| print("loaded fp32; nodes", len(model.graph.node), flush=True) | |
| # fp16: keep gather/index ops in fp32 via op_block_list defaults; blocks handle the int paths | |
| t0 = time.time() | |
| m16 = float16.convert_float_to_float16(model, keep_io_types=True, disable_shape_infer=True) | |
| onnx.save_model(m16, D + "/model_fp16.onnx", save_as_external_data=True, all_tensors_to_one_file=True, location="model_fp16.onnx.data", size_threshold=1024) | |
| print("fp16 done", round(time.time()-t0,1), "s", flush=True) | |
| del m16 | |
| # q4: MatMulNBits (what onnx-community ships as model_q4.onnx); embeddings stay fp32 | |
| t0 = time.time() | |
| q = MatMulNBitsQuantizer(onnx.load(D + "/model.onnx"), block_size=32, is_symmetric=True, accuracy_level=4) | |
| q.process() | |
| q.model.save_model_to_file(D + "/model_q4.onnx", use_external_data_format=True) | |
| print("q4 done", round(time.time()-t0,1), "s", flush=True) | |
| # q4f16: q4 matmuls with fp16 everything else | |
| t0 = time.time() | |
| m = onnx.load(D + "/model_q4.onnx") | |
| m = float16.convert_float_to_float16(m, keep_io_types=True, disable_shape_infer=True, op_block_list=float16.DEFAULT_OP_BLOCK_LIST + ["MatMulNBits"]) | |
| onnx.save_model(m, D + "/model_q4f16.onnx", save_as_external_data=True, all_tensors_to_one_file=True, location="model_q4f16.onnx.data", size_threshold=1024) | |
| print("q4f16 done", round(time.time()-t0,1), "s", flush=True) | |