Instructions to use vllm-sr/Vela-1.0-Encoder-307M-PII with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vllm-sr/Vela-1.0-Encoder-307M-PII with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="vllm-sr/Vela-1.0-Encoder-307M-PII")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-PII") model = AutoModelForTokenClassification.from_pretrained("vllm-sr/Vela-1.0-Encoder-307M-PII", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Unify Vela model cards with Semantic Router branding
Browse files- README.md +12 -7
- TECHNICAL.md +0 -163
README.md
CHANGED
|
@@ -19,14 +19,22 @@ tags:
|
|
| 19 |
- long-context
|
| 20 |
---
|
| 21 |
|
| 22 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
|
| 24 |
**Detect personal information before routing.**
|
| 25 |
|
| 26 |
307M · Multilingual · 32K context
|
| 27 |
|
| 28 |
-
[Collection](https://huggingface.co/collections/llm-semantic-router/vela-10-router-models-6aa555ba70cc6997d6d67798) · [vLLM Semantic Router](https://github.com/vllm-project/semantic-router)
|
| 29 |
-
|
| 30 |
Vela PII finds sensitive spans in multilingual text. Use its entity signals to support redaction, privacy-aware routing, and controlled data handling.
|
| 31 |
|
| 32 |
## What it brings
|
|
@@ -65,12 +73,9 @@ Choose the signals your router needs. Every model has a focused role.
|
|
| 65 |
| Role | Models |
|
| 66 |
|---|---|
|
| 67 |
| Understand requests | [Domain](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Domain) · [Feedback](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Feedback) · [Modality](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Modality) |
|
| 68 |
-
| Detect risks |
|
| 69 |
| Decide when to verify | [FactCheck](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-FactCheck) |
|
| 70 |
| Retrieve context | [Embedding](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Embedding) · [Reranker](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Reranker) |
|
| 71 |
| Build new capabilities | [Encoder](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M) |
|
| 72 |
|
| 73 |
¹ Coming soon. Explore available models in the [Vela collection](https://huggingface.co/collections/llm-semantic-router/vela-10-router-models-6aa555ba70cc6997d6d67798).
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
[Documentation and evaluation](TECHNICAL.md)
|
|
|
|
| 19 |
- long-context
|
| 20 |
---
|
| 21 |
|
| 22 |
+
<div align="center">
|
| 23 |
+
<img src="https://vllm-sr.ai/img/vllm-sr-logo.social.png" alt="vLLM Semantic Router" width="560" />
|
| 24 |
+
<p>
|
| 25 |
+
<a href="https://vllm-sr.ai/"><strong>Docs</strong></a> |
|
| 26 |
+
<a href="https://vllm-sr.ai/blog/"><strong>Blog</strong></a> |
|
| 27 |
+
<a href="https://vllm-dev.slack.com/archives/C09CTGF8KCN"><strong>Slack</strong></a> |
|
| 28 |
+
<a href="https://github.com/vllm-project/semantic-router"><strong>GitHub</strong></a>
|
| 29 |
+
</p>
|
| 30 |
+
</div>
|
| 31 |
+
|
| 32 |
+
# Vela-1.0-Encoder-307M-PII
|
| 33 |
|
| 34 |
**Detect personal information before routing.**
|
| 35 |
|
| 36 |
307M · Multilingual · 32K context
|
| 37 |
|
|
|
|
|
|
|
| 38 |
Vela PII finds sensitive spans in multilingual text. Use its entity signals to support redaction, privacy-aware routing, and controlled data handling.
|
| 39 |
|
| 40 |
## What it brings
|
|
|
|
| 73 |
| Role | Models |
|
| 74 |
|---|---|
|
| 75 |
| Understand requests | [Domain](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Domain) · [Feedback](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Feedback) · [Modality](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Modality) |
|
| 76 |
+
| Detect risks | Guard¹ · [PII](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-PII) · Safety¹ · Hazard¹ |
|
| 77 |
| Decide when to verify | [FactCheck](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-FactCheck) |
|
| 78 |
| Retrieve context | [Embedding](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Embedding) · [Reranker](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M-Reranker) |
|
| 79 |
| Build new capabilities | [Encoder](https://huggingface.co/llm-semantic-router/Vela-1.0-Encoder-307M) |
|
| 80 |
|
| 81 |
¹ Coming soon. Explore available models in the [Vela collection](https://huggingface.co/collections/llm-semantic-router/vela-10-router-models-6aa555ba70cc6997d6d67798).
|
|
|
|
|
|
|
|
|
TECHNICAL.md
DELETED
|
@@ -1,163 +0,0 @@
|
|
| 1 |
-
---
|
| 2 |
-
license: mit
|
| 3 |
-
library_name: transformers
|
| 4 |
-
pipeline_tag: token-classification
|
| 5 |
-
base_model: llm-semantic-router/mmbert-32k-yarn
|
| 6 |
-
language:
|
| 7 |
-
- en
|
| 8 |
-
- zh
|
| 9 |
-
- es
|
| 10 |
-
- fr
|
| 11 |
-
- de
|
| 12 |
-
- ja
|
| 13 |
-
tags:
|
| 14 |
-
- vela
|
| 15 |
-
- modernbert
|
| 16 |
-
- pii
|
| 17 |
-
- token-classification
|
| 18 |
-
- multilingual
|
| 19 |
-
- long-context
|
| 20 |
-
---
|
| 21 |
-
|
| 22 |
-
# Vela-1.0-Encoder-307M-PII
|
| 23 |
-
|
| 24 |
-
This is a repaired PII token classifier with the existing 35 BIO labels covering 17 entity types. On 1,704 frozen synthetic short requests, exact entity F1 increased from **0.201 to 0.936**. It also improves the authored 4K–32K stress sets. These results establish improvement on the documented synthetic test, not broad natural-document or production PII accuracy.
|
| 25 |
-
|
| 26 |
-
**Scanning remains valuable.** On 72 long inputs containing one sparse email needle, the final model recovered 55/72 complete emails with a full-sequence forward and 72/72 with overlapping 512-token windows. Capacity for 32K input does not make full-context inference the best default for this task. TITLE and NRP remain weak despite preserving their label indices.
|
| 27 |
-
|
| 28 |
-
## Architecture and actual parent
|
| 29 |
-
|
| 30 |
-
| Property | Value |
|
| 31 |
-
|---|---|
|
| 32 |
-
| Architecture | Standard `ModernBertForTokenClassification` |
|
| 33 |
-
| Shared encoder | 22 layers, width 768, 12 attention heads; 306,939,648 parameters |
|
| 34 |
-
| Total unique parameters, including prediction head and classifier | 307,557,155 |
|
| 35 |
-
| Output | 35 per-token BIO logits; no sequence pooling |
|
| 36 |
-
| Position capacity | 32,768 tokens including special tokens |
|
| 37 |
-
| Released weights | FP32 |
|
| 38 |
-
| Weight SHA-256 | `0317be9387f268cd9c959358e45f0750c097ebc64ad16ab6c2207f203405363c` |
|
| 39 |
-
|
| 40 |
-
This Vela task release **retains the original mmBERT encoder parent**, [mmbert-32k-yarn at revision 72a23a6640489471eb4ff7ad3ec5bc80af8a27de](https://huggingface.co/llm-semantic-router/mmbert-32k-yarn/tree/72a23a6640489471eb4ff7ad3ec5bc80af8a27de). It is not derived from the separately released Vela Base checkpoint. Vela names the versioned task family; a shared size and architecture do not establish identical weight lineage.
|
| 41 |
-
|
| 42 |
-
The repair continues [the original PII adapter at revision 58ecd71088283cd5ee660d5f2b155d32985a93b2](https://huggingface.co/llm-semantic-router/mmbert32k-pii-detector-lora/tree/58ecd71088283cd5ee660d5f2b155d32985a93b2). Before repair, its FP32 merge matched all 138 tensors of [the original merged baseline at revision d22c818cf9f2a8bfbbed8508cb417dc16a1ba3ea](https://huggingface.co/llm-semantic-router/mmbert32k-pii-detector-merged/tree/d22c818cf9f2a8bfbbed8508cb417dc16a1ba3ea) exactly. The new adapter was selected at stage 4, step 300 and frozen before final predictions. Its SHA-256 is `c7d1d1a301418c47a71aff87cb453b53cbd42083270f549c52abe31fa07ff461`; final FP32 merge checking observed a maximum logit difference of 4.39e-5.
|
| 43 |
-
|
| 44 |
-
The complete label order is stored in `config.json`, `pii_mapping.json`, and `label_mapping.json`, hashed together with the weights. Entity types are AGE, CREDIT_CARD, DATE_TIME, DOMAIN_NAME, EMAIL_ADDRESS, GPE, IBAN_CODE, IP_ADDRESS, NRP, ORGANIZATION, PERSON, PHONE_NUMBER, STREET_ADDRESS, TITLE, US_DRIVER_LICENSE, US_SSN, and ZIP_CODE. Preserving these output classes is not a claim that all types are reliable.
|
| 45 |
-
|
| 46 |
-
## Frozen FP32 final results
|
| 47 |
-
|
| 48 |
-
All four runs use the same 1,824 records, the same gold spans, FP32 parameters and forward computation, and the same permissive BIO decoder. Each predicted entity must match both type and exact trimmed boundaries. Orphan `I-*` starts an entity; different types are never joined to repair a mistake. No confidence threshold is applied. The original model is evaluated with the same repaired decoder, so the weight improvement is not a comparison against an older decoder bug.
|
| 49 |
-
|
| 50 |
-
| Final subset | Rows | Original full F1 | Vela full F1 | Original 512/64-window F1 | Vela 512/64-window F1 |
|
| 51 |
-
|---|---:|---:|---:|---:|---:|
|
| 52 |
-
| Synthetic short requests | 1,704 | 0.201 | 0.936 | 0.201 | 0.936 |
|
| 53 |
-
| Authored 4,096-token stress | 30 | 0.022 | 0.866 | 0.046 | 0.823 |
|
| 54 |
-
| Authored 8,192-token stress | 30 | 0.025 | 0.809 | 0.046 | 0.851 |
|
| 55 |
-
| Authored 16,384-token stress | 30 | 0.021 | 0.781 | 0.048 | 0.837 |
|
| 56 |
-
| Authored 32,768-token stress | 30 | 0.014 | 0.719 | 0.043 | 0.851 |
|
| 57 |
-
|
| 58 |
-
Each long bucket contains 18 sparse email needles, six packed synthetic records, and six negative inputs across six languages. Length/position variants reuse source payloads, and packed records repeat synthetic examples. They are correlated stress variants, not independent natural long documents. All gold spans remain in scoring, including boundaries that the fixed tokenizer cannot represent exactly.
|
| 59 |
-
|
| 60 |
-
| Complete-email recall subset | Original full | Vela full | Original windows | Vela windows |
|
| 61 |
-
|---|---:|---:|---:|---:|
|
| 62 |
-
| Short requests | 30/288 | 287/288 | 30/288 | 287/288 |
|
| 63 |
-
| Sparse long needles | 2/72 | 55/72 | 8/72 | 72/72 |
|
| 64 |
-
| Packed long records | 0/2,855 | 2,854/2,855 | 0/2,855 | 2,844/2,855 |
|
| 65 |
-
|
| 66 |
-
The many packed emails dominate aggregate email recall; that number would hide the full-sequence model's 17 missed sparse needles. Windowed inference recovers every sparse needle in this test but makes some additional packed boundary errors. The 120 negative documents, including 96 short and 24 long negatives, had no predicted entities with either Vela mode, compared with 11 original-model full-mode and four original-model window-mode false-positive documents.
|
| 67 |
-
|
| 68 |
-
Windowing is not uniformly better. Full mode has higher F1 on the 4K bucket; packed-context disambiguation and boundary fragments affect individual types and languages. Overall mixed synthetic entity F1 is 0.771 full versus 0.853 windows, but packed entities dominate this aggregate. Report the rows above instead of treating that mixed score as natural-request accuracy.
|
| 69 |
-
|
| 70 |
-
### Unresolved type and language limitations
|
| 71 |
-
|
| 72 |
-
On short inputs, TITLE has F1 **0.000** with 12 gold entities, NRP **0.273** with 24, and GPE **0.646** with 36. The old model also scored zero on TITLE. Across the mixed test, TITLE remains 0.014 full and zero windowed; NRP is 0.087 full and 0.375 windowed. These labels should not be advertised as production-ready solely because they exist in the model configuration.
|
| 73 |
-
|
| 74 |
-
All six languages have higher measured entity F1 than the original model, but the language tests share narrow template and entity-value families. For example, mixed full F1 is 0.616 in French and 0.707 in German versus 0.844 in English; these figures are dominated by packed synthetic entities. Chinese full-mode F1 exceeds its window-mode F1 in this particular mixed test; the Japanese result favors windows. Broad multilingual robustness remains unestablished. Full per-type, language, kind, length, and position counts are retained in the linked reports.
|
| 75 |
-
|
| 76 |
-
## Why the repair was needed
|
| 77 |
-
|
| 78 |
-
The pinned Presidio training-source audit found 1,500 rows but only 1,392 unique texts, 49 email entities, and a maximum input length of 110 tokens. Those emails did not cover local-part punctuation such as dots, plus signs, underscores, or hyphens. The 49 `@` and domain-dot tokens had valid `I-EMAIL_ADDRESS` supervision; the audit found no email/domain overlapping spans or invalid BIO transitions. The evidence supports a coverage gap, not a claim that those punctuation labels were incorrectly assigned.
|
| 79 |
-
|
| 80 |
-
A direct random 80/20 split produced 14 duplicate texts across partitions, and every validation template family also appeared in training. The new repair corpus instead isolates generated template families, texts, source groups, and complete entity values across splits. The deduplicated Presidio rows are explicitly **seen retention**, since the original weights may already have trained on them. They are not the independent final set.
|
| 81 |
-
|
| 82 |
-
Alignment supervises every entity subword and rejects ambiguous overlapping spans. Four Presidio training rows whose exact span boundary crossed a token containing additional non-whitespace context were explicitly rejected. A separate development boundary audit found additional unrepresentable synthetic phone-number spans; these evaluation labels were preserved, not removed to improve scores.
|
| 83 |
-
|
| 84 |
-
## Training and rejected Base migration
|
| 85 |
-
|
| 86 |
-
The repair trains rank-32 LoRA adapters in attention/MLP projections and the existing token classifier, with alpha 64 and dropout 0.1. The original shared prediction head remains frozen. Training uses FP32 parameters, BF16 autocast, SDPA, non-reentrant checkpointing, finite-gradient checks, and explicitly normalized token-classification loss. The selected curriculum is:
|
| 87 |
-
|
| 88 |
-
| Stage | Maximum tokens | Steps run / selected | Microbatch / accumulation | Learning rate | Short/seen replay probability |
|
| 89 |
-
|---|---:|---:|---:|---:|---:|
|
| 90 |
-
| Short repair | 2,048 | 400 / 300 | 8 / 2 | 3e-5 | 0.40 |
|
| 91 |
-
| 4K continuation | 4,096 | 400 / 300 | 1 / 4 | 1e-5 | 0.65 |
|
| 92 |
-
| 8K continuation | 8,192 | 400 / 300 | 1 / 4 | 8e-6 | 0.65 |
|
| 93 |
-
| 32K continuation | 32,768 | 600 / 300 | 1 / 4 | 5e-6 | 0.65 |
|
| 94 |
-
|
| 95 |
-
The long stages balance replay by entity type and select using equally weighted length-bucket development F1. The short stage selects strict entity micro F1. The maximum is a resource ceiling; stage receipts record the actual sequences and sampled budgets. Rejected earlier curriculum attempts were not the selected lineage.
|
| 96 |
-
|
| 97 |
-
A controlled attempt to use Vela Base retained the original adapter/classifier save scope. Development 32K full F1 fell from 0.6202 to 0.6067, exceeding the prewritten maximum drop of 0.01. A bounded 200-step, 2e-6 adaptation did not recover the gate; the best development-selected checkpoint remained its initial state. The original-parent candidate was therefore retained without relaxing the gate or selecting on final results.
|
| 98 |
-
|
| 99 |
-
This migration also replaced the base's shared prediction-head parameters, because the original PII adapter does not save `head.*`. A CPU tensor audit verified changes in `head.dense.weight` and `head.norm.weight` as well as the Base change. Thus this is not an isolated causal test of encoder-only migration. [Migration evidence](evaluation/migration-decision.json) and [head-scope audit](evaluation/base-head-transfer-audit.json) preserve that distinction. No claim that Vela Base is generally worse for PII follows from this experiment.
|
| 100 |
-
|
| 101 |
-
## Data and licensing scope
|
| 102 |
-
|
| 103 |
-
**This repair phase** uses the [MIT-licensed Presidio Research synthetic source at revision f3ff907eba57b8d380711ce7ca82a42696cd0490](https://github.com/microsoft/presidio-research/tree/f3ff907eba57b8d380711ce7ca82a42696cd0490) and new authored synthetic data. The source file SHA-256 is `ec08a771ba8135314cafb60752b2295212222ba3a4cd75d73811839c699e0012`. No AI4Privacy dataset was downloaded for the repair. The inherited original PII adapter's MIT-licensed card lists both AI4Privacy and Presidio and lacks a complete historical row manifest; the current data audit does not prove its entire training history used only the repair sources.
|
| 104 |
-
|
| 105 |
-
The new corpus has 6,372 training, 888 development, 1,704 final short rows, and 1,392 separate seen-retention rows. The long generator adds 30 records per split per length. All four short files and all 12 long files were independently regenerated and matched the recorded bytes. The generators, immutable source references, stage receipts, and data hashes are included under `reproducibility/`. Model weights use MIT; copied Semantic Router training code uses Apache-2.0; source attribution is retained in `NOTICE.md`.
|
| 106 |
-
|
| 107 |
-
## Runtime scope and measurement
|
| 108 |
-
|
| 109 |
-
HF token-window evaluation uses a 512-token maximum with a 64-token overlap, global offsets, and deduplication only when entity type and boundaries are identical. It is **not the Go router's rune-budget scanning algorithm**, and these results do not include serving-policy confidence thresholds. Reported offsets use Unicode codepoints; native interfaces may use UTF-8 bytes and need explicit conversion.
|
| 110 |
-
|
| 111 |
-
Across the fixed final batch, Vela full mode ran 1,824 sequences in 234 forwards with 177.83 seconds of measured GPU forward time. Window mode ran 5,904 windows in 369 forwards with 17.62 seconds of measured forward time. These are summed forward measurements under the recorded evaluation conditions, excluding tokenization, HTTP, scheduler, and queueing overhead. They are not per-request latency, isolated throughput, or an end-to-end serving speedup guarantee.
|
| 112 |
-
|
| 113 |
-
The release uses FP32 weights. The portable FP32 ONNX export passed the original tight probability tolerance and exact BIO-label comparison on the recorded short/padded probes. The experimental FP16 export failed at batch 2 × 512 tokens (maximum probability difference 0.00870889); an FP32 LayerNorm-only diagnostic did not consistently fix it. FP16 is not a qualified release artifact. These numerical checks establish export equivalence on their probe set, not PII task accuracy.
|
| 114 |
-
|
| 115 |
-
A 32K acceptance limit is distinct from dependable detection, suitable memory budgets, and an engine's tested support. Actual native CPU/GPU entity spans, confidence thresholds, UTF-8 byte offsets, and long-input/overflow checks require separate reports; none are inferred from the HF task-quality table above.
|
| 116 |
-
|
| 117 |
-
See [frozen candidate hashes](evaluation/candidate-lock.json), [full FP32 results](evaluation/final-candidate-full-fp32.json), [window FP32 results](evaluation/final-candidate-window-fp32.json), and [reproduction instructions](reproducibility/README.md). The final weights and test predictions were not changed to remove the documented failures.
|
| 118 |
-
|
| 119 |
-
## Qualified native and ONNX variants
|
| 120 |
-
|
| 121 |
-
The released `onnx/model.onnx` uses FP32 query-block attention and fixes both
|
| 122 |
-
input batch dimensions to **one request**. Batches larger than one are explicitly
|
| 123 |
-
rejected. This matches the Router's single-request classifier interface; it is
|
| 124 |
-
not a claim that all dynamic batching modes passed. The adjacent external weight
|
| 125 |
-
file is required. No CK custom-op library is needed for this variant.
|
| 126 |
-
|
| 127 |
-
On the recorded AMD ROCm execution provider, native interface comparisons cover
|
| 128 |
-
short inputs and 4K, 8K, 16K and 32K without truncation; inputs above 32,768 tokens
|
| 129 |
-
are rejected. The largest native probability/confidence difference was
|
| 130 |
-
0.00001276. The ONNX-level report separately checks the complete output
|
| 131 |
-
probability vector, unchanged argmax labels, and GPU execution of attention
|
| 132 |
-
inside each loop. See [native AMD evidence](evaluation/engines/vela-pii-blocked-b1-native-gpu-32k-v1.json)
|
| 133 |
-
and [graph/profile evidence](evaluation/engines/vela-pii-blocked-b1-onnx-gpu-v1.json).
|
| 134 |
-
Candle CPU evidence also covers all 24 native probes through 32K, including
|
| 135 |
-
UTF-8 entity spans and overflow rejection (maximum confidence difference
|
| 136 |
-
0.00008213). The single 32K forward took approximately 34 minutes with eight
|
| 137 |
-
CPU threads under the recorded shared-host conditions, demonstrating numerical
|
| 138 |
-
support rather than practical online latency; long-context HF task-quality tables do not substitute for native
|
| 139 |
-
engine qualification. These are equivalence tests on fixed inputs, not another
|
| 140 |
-
independent accuracy test or a serving-latency benchmark.
|
| 141 |
-
|
| 142 |
-
The `lora/` variant retains the exact development-selected adapter and pinned
|
| 143 |
-
parent. Use its `load_adapter.py` example to restore task configuration and
|
| 144 |
-
pooling before loading weights. Independent padded bilingual/Unicode CPU
|
| 145 |
-
comparisons against the frozen merged checkpoint are recorded in
|
| 146 |
-
`standalone-verification.json`; the adapter does not silently change lineage.
|
| 147 |
-
See [ONNX reproduction](reproducibility/onnx/README.md).
|
| 148 |
-
|
| 149 |
-
The experimental dynamic-B2 8K graph did not pass the strict full-probability
|
| 150 |
-
gate: maximum difference 0.00078665, with no BIO argmax change. On exactly the
|
| 151 |
-
same input, the unmodified portable FP32 export differed from HF by 0.00061213;
|
| 152 |
-
the two ONNX variants differed by 0.00017452. Other B2 long probes passed, but
|
| 153 |
-
these do not establish general B2 qualification. These failed diagnostic files
|
| 154 |
-
are retained under `evaluation/engines/`; the experimental dynamic graph and
|
| 155 |
-
failed FP16/CK variants are excluded from the release. The B1 artifact passed
|
| 156 |
-
all twelve recorded lengths without changing the numerical gate. Native PII
|
| 157 |
-
checks additionally compare entity text, types, UTF-8 byte spans and confidence,
|
| 158 |
-
and reject null, invalid UTF-8 and overflowing inputs.
|
| 159 |
-
|
| 160 |
-
The [storage provenance](evaluation/engines/blocked-storage-provenance.json) binds
|
| 161 |
-
the released graph and external data file to the qualified graph and the frozen
|
| 162 |
-
FP32 export. All 157 retained initializer tensors match their original dtype,
|
| 163 |
-
shape and bytes; graph rewriting did not alter task weights.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|