nico-martin HF Staff commited on
Commit
3bc2553
·
verified ·
1 Parent(s): f70887c

Model card: minimal usage section, original card marked, minimal tags

Browse files
Files changed (1) hide show
  1. README.md +24 -84
README.md CHANGED
@@ -4,78 +4,23 @@ base_model: microsoft/deberta-v3-large
4
  pipeline_tag: text-classification
5
  library_name: transformers.js
6
  tags:
7
- - typed-decisions
8
- - calibrated
9
- - decision-model
10
- - open-jev
11
- - deberta-v3
12
- - text-classification
13
  - deberta-v2
 
14
  - onnx
15
  - transformers.js
16
- - transformersjs
17
  ---
18
 
19
  > This is a Transformers.js-ready ONNX conversion of the original Hugging Face model [com-kotobalabs/open-jev-deberta-v3-large](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large).
20
  > The original model card content follows below.
21
 
22
- ## ONNX conversion notes
23
-
24
- One graph runs the DeBERTa-v3-large encoder **and** the span-matching head, so
25
- a state with any number of typed questions is answered in a single forward
26
- pass, exactly like the source `OpenJev.decide()`.
27
-
28
- ### Graph interface
29
-
30
- | name | shape | dtype | meaning |
31
- | ---------------- | ------------------------- | ----- | ------- |
32
- | `input_ids` | `[batch, seq]` | int64 | `[CLS] [STATE] state [Q] instr [OPT] opt … [SEP]`, right padded with `[PAD]` (0) |
33
- | `attention_mask` | `[batch, seq]` | int64 | 1 for real tokens |
34
- | `seg` | `[batch, seq]` | int64 | span slot of each token: option text tokens get their pair index, question text tokens get `pairs + question index`, everything else `-1` |
35
- | `pair_q` | `[batch, pairs]` | int64 | slot of the question of every (question, option) pair |
36
- | `pair_opt` | `[batch, pairs]` | int64 | slot of the option of every pair |
37
- | `logits` | `[batch, pairs]` | float32 (float16 in `fp16`/`q4f16`) | raw score per pair |
38
-
39
- Divide the logits of one question's pairs by the temperature (1.05) and
40
- softmax them; `choice` = argmax, `score` = expected level index, yes/no
41
- questions use the fixed option list `["no", "yes"]` and report p(yes). Batching
42
- several states: right-pad `input_ids`/`seg` (pad `seg` with `-1`) and pad the
43
- pair tensors with an unused slot such as `-2`.
44
-
45
- Pooling is done with equality masks and matrix products, so batch size,
46
- sequence length, and pair count are all dynamic and the graph has no
47
- data-dependent shapes. Context is 512 tokens (state cut to 256), as in the
48
- source. The three marker tokens (`[STATE]` 128001, `[Q]` 128002, `[OPT]`
49
- 128003) are in `tokenizer.json`.
50
-
51
- ### Variants
52
-
53
- | file | weights | size |
54
- | --- | --- | --- |
55
- | `onnx/model.onnx` | fp32 | 1.75 GB |
56
- | `onnx/model_fp16.onnx` | fp16, position-bucket math kept fp32 | 0.87 GB |
57
- | `onnx/model_q4.onnx` | 4-bit `MatMulNBits` (block 32, `accuracy_level=4`) + 4-bit embedding, fp32 activations | 0.48 GB |
58
- | `onnx/model_q4f16.onnx` | same 4-bit weights, fp16 activations | 0.34 GB |
59
-
60
- Accuracy against the original PyTorch model on 15 typed questions (banking77,
61
- SST-5, BoolQ rows and the example above): `fp32` and `fp16` reproduce every
62
- answer (max logit difference 1e-5 and 0.008); `q4` and `q4f16` agree on 14 of
63
- 15 answers with probability shifts up to 0.2. No `q8` (`model_quantized.onnx`)
64
- is published: dynamic uint8 quantization, per-tensor or per-channel, changed
65
- the answer on 8 of the 15 questions, so it was rejected. Pass `dtype: "fp16"`
66
- (default), `"q4f16"`, `"q4"`, or `"fp32"`.
67
-
68
- Speed on an Apple M3 Max with WebGPU: the example below (3 questions, 17
69
- options, 103 tokens) takes about 50-65 ms per pass at `fp16` or `q4f16` once
70
- the model is warm; `fp32` on CPU in Node takes about 130 ms.
71
-
72
- ### Transformers.js
73
-
74
- Load with `AutoModel` (the config is `deberta-v2`, so Transformers.js picks
75
- `DebertaV2Model`; its encoder forward feeds every graph input by name) and
76
- `AutoTokenizer` (`DebertaV2Tokenizer` reads the marker tokens from
77
- `tokenizer.json`). Everything below runs in the browser (`device: "webgpu"` or
78
- WASM) and in Node.
79
 
80
  ```js
81
  import { AutoModel, AutoTokenizer, Tensor } from "@huggingface/transformers";
@@ -87,25 +32,29 @@ const model = await AutoModel.from_pretrained(repo, { dtype: "fp16", device: "we
87
  const enc = (t) => Array.from(tokenizer(t, { add_special_tokens: false }).input_ids.data, Number);
88
  const [CLS, SEP, STATE, Q, OPT] = ["[CLS]", "[SEP]", "[STATE]", "[Q]", "[OPT]"].map((m) => enc(m)[0]);
89
 
90
- // one state, a list of questions ({type, instructions, options}); noul = ["no","yes"]
 
 
 
 
 
91
  const tokens = [CLS, STATE, ...enc(state).slice(0, 256)];
92
- const seg = new Array(tokens.length).fill(-1);
93
  const pairQ = [], pairOpt = [], groups = [];
94
  const totalPairs = questions.reduce((n, q) => n + q.options.length, 0);
95
  questions.forEach((q, qi) => {
96
- const qSlot = totalPairs + qi;
97
  const text = enc(q.instructions);
98
- tokens.push(Q, ...text); seg.push(-1, ...text.map(() => qSlot));
99
  groups.push(q.options.map((option) => {
100
  const optText = enc(option);
101
  tokens.push(OPT, ...optText); seg.push(-1, ...optText.map(() => pairOpt.length));
102
- pairQ.push(qSlot); pairOpt.push(pairOpt.length);
103
  return pairOpt.length - 1;
104
  }));
105
  });
106
  tokens.push(SEP); seg.push(-1);
107
 
108
- const i64 = (values, dims) => new Tensor("int64", BigInt64Array.from(values, BigInt), dims);
109
  const { logits } = await model({
110
  input_ids: i64(tokens, [1, tokens.length]),
111
  attention_mask: i64(tokens.map(() => 1), [1, tokens.length]),
@@ -116,24 +65,15 @@ const { logits } = await model({
116
  const scores = Array.from(logits.to("float32").data);
117
  const softmax = (xs) => { const m = Math.max(...xs); const e = xs.map((x) => Math.exp((x - m) / 1.05)); const s = e.reduce((a, b) => a + b); return e.map((x) => x / s); };
118
  const answers = groups.map((g) => softmax(g.map((p) => scores[p])));
119
- // choice: argmax; score: expected level index; noul: probability of "yes" (index 1)
120
  ```
121
 
122
- To batch several states, right-pad `input_ids` with `0`, `attention_mask` with
123
- `0`, `seg` with `-1`, and the pair tensors with an unused slot such as `-2`;
124
- padded pairs produce meaningless logits that you simply ignore. In Node the
125
- same code runs with `device: "cpu"` or `"webgpu"`.
126
-
127
- ### Limitations of the conversion
128
 
129
- - Standard ONNX operators only (opset 17); runs on any ONNX Runtime including
130
- onnxruntime-web WASM and WebGPU.
131
- - The graph answers questions that fit the 512-token window with the state;
132
- the host must split larger question sets across calls (the source raises
133
- instead).
134
- - Everything else (English only, OOD gap, calibration) is inherited from the
135
- source model and described below.
136
 
 
137
 
138
  # open-jev-deberta-v3-large
139
 
 
4
  pipeline_tag: text-classification
5
  library_name: transformers.js
6
  tags:
 
 
 
 
 
 
7
  - deberta-v2
8
+ - open-jev
9
  - onnx
10
  - transformers.js
 
11
  ---
12
 
13
  > This is a Transformers.js-ready ONNX conversion of the original Hugging Face model [com-kotobalabs/open-jev-deberta-v3-large](https://huggingface.co/com-kotobalabs/open-jev-deberta-v3-large).
14
  > The original model card content follows below.
15
 
16
+ ## Usage with Transformers.js
17
+
18
+ Variants under `onnx/`: `fp16` (default, 0.88 GB), `fp32` (1.75 GB), `q4`
19
+ (0.48 GB), `q4f16` (0.35 GB). One state plus any number of typed questions go
20
+ into a single sequence; one forward pass returns a logit per (question, option)
21
+ pair. The graph takes the token ids, an attention mask, a span slot per token
22
+ (`seg`: option tokens get their pair index, question tokens get
23
+ `pairs + question index`, everything else `-1`) and the slot ids of each pair.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24
 
25
  ```js
26
  import { AutoModel, AutoTokenizer, Tensor } from "@huggingface/transformers";
 
32
  const enc = (t) => Array.from(tokenizer(t, { add_special_tokens: false }).input_ids.data, Number);
33
  const [CLS, SEP, STATE, Q, OPT] = ["[CLS]", "[SEP]", "[STATE]", "[Q]", "[OPT]"].map((m) => enc(m)[0]);
34
 
35
+ const state = "I was charged twice for the same order. I want my money back now.";
36
+ const questions = [
37
+ { type: "choice", instructions: "Which product area is the message about?", options: ["fees & charges", "refund & dispute", "card", "other"] },
38
+ { type: "noul", instructions: "The customer is asking for a refund.", options: ["no", "yes"] },
39
+ ];
40
+
41
  const tokens = [CLS, STATE, ...enc(state).slice(0, 256)];
42
+ const seg = tokens.map(() => -1);
43
  const pairQ = [], pairOpt = [], groups = [];
44
  const totalPairs = questions.reduce((n, q) => n + q.options.length, 0);
45
  questions.forEach((q, qi) => {
 
46
  const text = enc(q.instructions);
47
+ tokens.push(Q, ...text); seg.push(-1, ...text.map(() => totalPairs + qi));
48
  groups.push(q.options.map((option) => {
49
  const optText = enc(option);
50
  tokens.push(OPT, ...optText); seg.push(-1, ...optText.map(() => pairOpt.length));
51
+ pairQ.push(totalPairs + qi); pairOpt.push(pairOpt.length);
52
  return pairOpt.length - 1;
53
  }));
54
  });
55
  tokens.push(SEP); seg.push(-1);
56
 
57
+ const i64 = (v, dims) => new Tensor("int64", BigInt64Array.from(v, BigInt), dims);
58
  const { logits } = await model({
59
  input_ids: i64(tokens, [1, tokens.length]),
60
  attention_mask: i64(tokens.map(() => 1), [1, tokens.length]),
 
65
  const scores = Array.from(logits.to("float32").data);
66
  const softmax = (xs) => { const m = Math.max(...xs); const e = xs.map((x) => Math.exp((x - m) / 1.05)); const s = e.reduce((a, b) => a + b); return e.map((x) => x / s); };
67
  const answers = groups.map((g) => softmax(g.map((p) => scores[p])));
68
+ // choice: argmax; score: expected level index; noul (options ["no","yes"]): answers[i][1] is p(yes)
69
  ```
70
 
71
+ The sequence is limited to 512 tokens (state cut to 256). Temperature 1.05 is
72
+ the source model's calibration.
 
 
 
 
73
 
74
+ ---
 
 
 
 
 
 
75
 
76
+ ## Original model card
77
 
78
  # open-jev-deberta-v3-large
79