Instructions to use onnx-community/open-jev-deberta-v3-large-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use onnx-community/open-jev-deberta-v3-large-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-classification', 'onnx-community/open-jev-deberta-v3-large-ONNX');
Model card: minimal usage section, original card marked, minimal tags
Browse files
README.md
CHANGED
|
@@ -4,78 +4,23 @@ base_model: microsoft/deberta-v3-large
|
|
| 4 |
pipeline_tag: text-classification
|
| 5 |
library_name: transformers.js
|
| 6 |
tags:
|
| 7 |
-
- typed-decisions
|
| 8 |
-
- calibrated
|
| 9 |
-
- decision-model
|
| 10 |
-
- open-jev
|
| 11 |
-
- deberta-v3
|
| 12 |
-
- text-classification
|
| 13 |
- deberta-v2
|
|
|
|
| 14 |
- onnx
|
| 15 |
- transformers.js
|
| 16 |
-
- transformersjs
|
| 17 |
---
|
| 18 |
|
| 19 |
> This is a Transformers.js-ready ONNX conversion of the original Hugging Face model [com-kotobalabs/open-jev-deberta-v3-large](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large).
|
| 20 |
> The original model card content follows below.
|
| 21 |
|
| 22 |
-
##
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
pass
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
| name | shape | dtype | meaning |
|
| 31 |
-
| ---------------- | ------------------------- | ----- | ------- |
|
| 32 |
-
| `input_ids` | `[batch, seq]` | int64 | `[CLS] [STATE] state [Q] instr [OPT] opt … [SEP]`, right padded with `[PAD]` (0) |
|
| 33 |
-
| `attention_mask` | `[batch, seq]` | int64 | 1 for real tokens |
|
| 34 |
-
| `seg` | `[batch, seq]` | int64 | span slot of each token: option text tokens get their pair index, question text tokens get `pairs + question index`, everything else `-1` |
|
| 35 |
-
| `pair_q` | `[batch, pairs]` | int64 | slot of the question of every (question, option) pair |
|
| 36 |
-
| `pair_opt` | `[batch, pairs]` | int64 | slot of the option of every pair |
|
| 37 |
-
| `logits` | `[batch, pairs]` | float32 (float16 in `fp16`/`q4f16`) | raw score per pair |
|
| 38 |
-
|
| 39 |
-
Divide the logits of one question's pairs by the temperature (1.05) and
|
| 40 |
-
softmax them; `choice` = argmax, `score` = expected level index, yes/no
|
| 41 |
-
questions use the fixed option list `["no", "yes"]` and report p(yes). Batching
|
| 42 |
-
several states: right-pad `input_ids`/`seg` (pad `seg` with `-1`) and pad the
|
| 43 |
-
pair tensors with an unused slot such as `-2`.
|
| 44 |
-
|
| 45 |
-
Pooling is done with equality masks and matrix products, so batch size,
|
| 46 |
-
sequence length, and pair count are all dynamic and the graph has no
|
| 47 |
-
data-dependent shapes. Context is 512 tokens (state cut to 256), as in the
|
| 48 |
-
source. The three marker tokens (`[STATE]` 128001, `[Q]` 128002, `[OPT]`
|
| 49 |
-
128003) are in `tokenizer.json`.
|
| 50 |
-
|
| 51 |
-
### Variants
|
| 52 |
-
|
| 53 |
-
| file | weights | size |
|
| 54 |
-
| --- | --- | --- |
|
| 55 |
-
| `onnx/model.onnx` | fp32 | 1.75 GB |
|
| 56 |
-
| `onnx/model_fp16.onnx` | fp16, position-bucket math kept fp32 | 0.87 GB |
|
| 57 |
-
| `onnx/model_q4.onnx` | 4-bit `MatMulNBits` (block 32, `accuracy_level=4`) + 4-bit embedding, fp32 activations | 0.48 GB |
|
| 58 |
-
| `onnx/model_q4f16.onnx` | same 4-bit weights, fp16 activations | 0.34 GB |
|
| 59 |
-
|
| 60 |
-
Accuracy against the original PyTorch model on 15 typed questions (banking77,
|
| 61 |
-
SST-5, BoolQ rows and the example above): `fp32` and `fp16` reproduce every
|
| 62 |
-
answer (max logit difference 1e-5 and 0.008); `q4` and `q4f16` agree on 14 of
|
| 63 |
-
15 answers with probability shifts up to 0.2. No `q8` (`model_quantized.onnx`)
|
| 64 |
-
is published: dynamic uint8 quantization, per-tensor or per-channel, changed
|
| 65 |
-
the answer on 8 of the 15 questions, so it was rejected. Pass `dtype: "fp16"`
|
| 66 |
-
(default), `"q4f16"`, `"q4"`, or `"fp32"`.
|
| 67 |
-
|
| 68 |
-
Speed on an Apple M3 Max with WebGPU: the example below (3 questions, 17
|
| 69 |
-
options, 103 tokens) takes about 50-65 ms per pass at `fp16` or `q4f16` once
|
| 70 |
-
the model is warm; `fp32` on CPU in Node takes about 130 ms.
|
| 71 |
-
|
| 72 |
-
### Transformers.js
|
| 73 |
-
|
| 74 |
-
Load with `AutoModel` (the config is `deberta-v2`, so Transformers.js picks
|
| 75 |
-
`DebertaV2Model`; its encoder forward feeds every graph input by name) and
|
| 76 |
-
`AutoTokenizer` (`DebertaV2Tokenizer` reads the marker tokens from
|
| 77 |
-
`tokenizer.json`). Everything below runs in the browser (`device: "webgpu"` or
|
| 78 |
-
WASM) and in Node.
|
| 79 |
|
| 80 |
```js
|
| 81 |
import { AutoModel, AutoTokenizer, Tensor } from "@huggingface/transformers";
|
|
@@ -87,25 +32,29 @@ const model = await AutoModel.from_pretrained(repo, { dtype: "fp16", device: "we
|
|
| 87 |
const enc = (t) => Array.from(tokenizer(t, { add_special_tokens: false }).input_ids.data, Number);
|
| 88 |
const [CLS, SEP, STATE, Q, OPT] = ["[CLS]", "[SEP]", "[STATE]", "[Q]", "[OPT]"].map((m) => enc(m)[0]);
|
| 89 |
|
| 90 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
const tokens = [CLS, STATE, ...enc(state).slice(0, 256)];
|
| 92 |
-
const seg =
|
| 93 |
const pairQ = [], pairOpt = [], groups = [];
|
| 94 |
const totalPairs = questions.reduce((n, q) => n + q.options.length, 0);
|
| 95 |
questions.forEach((q, qi) => {
|
| 96 |
-
const qSlot = totalPairs + qi;
|
| 97 |
const text = enc(q.instructions);
|
| 98 |
-
tokens.push(Q, ...text); seg.push(-1, ...text.map(() =>
|
| 99 |
groups.push(q.options.map((option) => {
|
| 100 |
const optText = enc(option);
|
| 101 |
tokens.push(OPT, ...optText); seg.push(-1, ...optText.map(() => pairOpt.length));
|
| 102 |
-
pairQ.push(
|
| 103 |
return pairOpt.length - 1;
|
| 104 |
}));
|
| 105 |
});
|
| 106 |
tokens.push(SEP); seg.push(-1);
|
| 107 |
|
| 108 |
-
const i64 = (
|
| 109 |
const { logits } = await model({
|
| 110 |
input_ids: i64(tokens, [1, tokens.length]),
|
| 111 |
attention_mask: i64(tokens.map(() => 1), [1, tokens.length]),
|
|
@@ -116,24 +65,15 @@ const { logits } = await model({
|
|
| 116 |
const scores = Array.from(logits.to("float32").data);
|
| 117 |
const softmax = (xs) => { const m = Math.max(...xs); const e = xs.map((x) => Math.exp((x - m) / 1.05)); const s = e.reduce((a, b) => a + b); return e.map((x) => x / s); };
|
| 118 |
const answers = groups.map((g) => softmax(g.map((p) => scores[p])));
|
| 119 |
-
// choice: argmax; score: expected level index; noul
|
| 120 |
```
|
| 121 |
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
padded pairs produce meaningless logits that you simply ignore. In Node the
|
| 125 |
-
same code runs with `device: "cpu"` or `"webgpu"`.
|
| 126 |
-
|
| 127 |
-
### Limitations of the conversion
|
| 128 |
|
| 129 |
-
-
|
| 130 |
-
onnxruntime-web WASM and WebGPU.
|
| 131 |
-
- The graph answers questions that fit the 512-token window with the state;
|
| 132 |
-
the host must split larger question sets across calls (the source raises
|
| 133 |
-
instead).
|
| 134 |
-
- Everything else (English only, OOD gap, calibration) is inherited from the
|
| 135 |
-
source model and described below.
|
| 136 |
|
|
|
|
| 137 |
|
| 138 |
# open-jev-deberta-v3-large
|
| 139 |
|
|
|
|
| 4 |
pipeline_tag: text-classification
|
| 5 |
library_name: transformers.js
|
| 6 |
tags:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
- deberta-v2
|
| 8 |
+
- open-jev
|
| 9 |
- onnx
|
| 10 |
- transformers.js
|
|
|
|
| 11 |
---
|
| 12 |
|
| 13 |
> This is a Transformers.js-ready ONNX conversion of the original Hugging Face model [com-kotobalabs/open-jev-deberta-v3-large](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large).
|
| 14 |
> The original model card content follows below.
|
| 15 |
|
| 16 |
+
## Usage with Transformers.js
|
| 17 |
+
|
| 18 |
+
Variants under `onnx/`: `fp16` (default, 0.88 GB), `fp32` (1.75 GB), `q4`
|
| 19 |
+
(0.48 GB), `q4f16` (0.35 GB). One state plus any number of typed questions go
|
| 20 |
+
into a single sequence; one forward pass returns a logit per (question, option)
|
| 21 |
+
pair. The graph takes the token ids, an attention mask, a span slot per token
|
| 22 |
+
(`seg`: option tokens get their pair index, question tokens get
|
| 23 |
+
`pairs + question index`, everything else `-1`) and the slot ids of each pair.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
|
| 25 |
```js
|
| 26 |
import { AutoModel, AutoTokenizer, Tensor } from "@huggingface/transformers";
|
|
|
|
| 32 |
const enc = (t) => Array.from(tokenizer(t, { add_special_tokens: false }).input_ids.data, Number);
|
| 33 |
const [CLS, SEP, STATE, Q, OPT] = ["[CLS]", "[SEP]", "[STATE]", "[Q]", "[OPT]"].map((m) => enc(m)[0]);
|
| 34 |
|
| 35 |
+
const state = "I was charged twice for the same order. I want my money back now.";
|
| 36 |
+
const questions = [
|
| 37 |
+
{ type: "choice", instructions: "Which product area is the message about?", options: ["fees & charges", "refund & dispute", "card", "other"] },
|
| 38 |
+
{ type: "noul", instructions: "The customer is asking for a refund.", options: ["no", "yes"] },
|
| 39 |
+
];
|
| 40 |
+
|
| 41 |
const tokens = [CLS, STATE, ...enc(state).slice(0, 256)];
|
| 42 |
+
const seg = tokens.map(() => -1);
|
| 43 |
const pairQ = [], pairOpt = [], groups = [];
|
| 44 |
const totalPairs = questions.reduce((n, q) => n + q.options.length, 0);
|
| 45 |
questions.forEach((q, qi) => {
|
|
|
|
| 46 |
const text = enc(q.instructions);
|
| 47 |
+
tokens.push(Q, ...text); seg.push(-1, ...text.map(() => totalPairs + qi));
|
| 48 |
groups.push(q.options.map((option) => {
|
| 49 |
const optText = enc(option);
|
| 50 |
tokens.push(OPT, ...optText); seg.push(-1, ...optText.map(() => pairOpt.length));
|
| 51 |
+
pairQ.push(totalPairs + qi); pairOpt.push(pairOpt.length);
|
| 52 |
return pairOpt.length - 1;
|
| 53 |
}));
|
| 54 |
});
|
| 55 |
tokens.push(SEP); seg.push(-1);
|
| 56 |
|
| 57 |
+
const i64 = (v, dims) => new Tensor("int64", BigInt64Array.from(v, BigInt), dims);
|
| 58 |
const { logits } = await model({
|
| 59 |
input_ids: i64(tokens, [1, tokens.length]),
|
| 60 |
attention_mask: i64(tokens.map(() => 1), [1, tokens.length]),
|
|
|
|
| 65 |
const scores = Array.from(logits.to("float32").data);
|
| 66 |
const softmax = (xs) => { const m = Math.max(...xs); const e = xs.map((x) => Math.exp((x - m) / 1.05)); const s = e.reduce((a, b) => a + b); return e.map((x) => x / s); };
|
| 67 |
const answers = groups.map((g) => softmax(g.map((p) => scores[p])));
|
| 68 |
+
// choice: argmax; score: expected level index; noul (options ["no","yes"]): answers[i][1] is p(yes)
|
| 69 |
```
|
| 70 |
|
| 71 |
+
The sequence is limited to 512 tokens (state cut to 256). Temperature 1.05 is
|
| 72 |
+
the source model's calibration.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 73 |
|
| 74 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 75 |
|
| 76 |
+
## Original model card
|
| 77 |
|
| 78 |
# open-jev-deberta-v3-large
|
| 79 |
|