Publish embeddinggemma-300m-memory-ft-v2 (qualified ONNX export and complete attribution)
Browse files- .gitattributes +2 -0
- GEMMA_PROHIBITED_USE_POLICY.txt +39 -0
- LICENSE +130 -0
- MODIFICATIONS.md +49 -0
- NOTICE +11 -0
- README.md +117 -0
- added_tokens.json +3 -0
- config.json +61 -0
- evaluation/README.md +83 -0
- evaluation/metrics.py +219 -0
- evaluation/promotion-summary.json +1 -0
- evaluation/promotion.json +0 -0
- evaluation/serving-qualification.json +134 -0
- model.onnx +3 -0
- model.onnx.data +3 -0
- publication-manifest.json +114 -0
- serving.json +28 -0
- special_tokens_map.json +33 -0
- tokenizer.json +3 -0
- tokenizer_config.json +0 -0
- vulkan-derivation.json +27 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
model.onnx.data filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
GEMMA_PROHIBITED_USE_POLICY.txt
ADDED
|
@@ -0,0 +1,39 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Gemma Prohibited Use Policy
|
| 2 |
+
Source: https://ai.google.dev/gemma/prohibited_use_policy
|
| 3 |
+
Retrieved: 2026-09-01
|
| 4 |
+
|
| 5 |
+
Google reserves the right to update this Gemma Prohibited Use Policy from time to time.
|
| 6 |
+
|
| 7 |
+
Last modified: February 21, 2024
|
| 8 |
+
|
| 9 |
+
You may not use nor allow others to use Gemma or Model Derivatives to:
|
| 10 |
+
|
| 11 |
+
1. Generate any content, including the outputs or results generated by Gemma or Model Derivatives, that infringes, misappropriates, or otherwise violates any individual's or entity's rights (including, but not limited to rights in copyrighted content).
|
| 12 |
+
2. Perform or facilitate dangerous, illegal, or malicious activities, including:
|
| 13 |
+
1. Facilitation or promotion of illegal activities or violations of law, such as:
|
| 14 |
+
1. Promoting or generating content related to child sexual abuse or exploitation;
|
| 15 |
+
2. Promoting or facilitating sale of, or providing instructions for synthesizing or accessing, illegal substances, goods, or services;
|
| 16 |
+
3. Facilitating or encouraging users to commit any type of crimes; or
|
| 17 |
+
4. Promoting or generating violent extremism or terrorist content.
|
| 18 |
+
2. Engagement in the illegal or unlicensed practice of any vocation or profession including, but not limited to, legal, medical, accounting, or financial professional practices.
|
| 19 |
+
3. Abuse, harm, interference, or disruption of services (or enable others to do the same), such as:
|
| 20 |
+
1. Promoting or facilitating the generation or distribution of spam; or
|
| 21 |
+
2. Generating content for deceptive or fraudulent activities, scams, phishing, or malware.
|
| 22 |
+
4. Attempts to override or circumvent safety filters or intentionally drive Gemma or Model Derivatives to act in a manner that contravenes this Gemma Prohibited Use Policy.
|
| 23 |
+
5. Generation of content that may harm or promote the harm of individuals or a group, such as:
|
| 24 |
+
1. Generating content that promotes or encourages hatred;
|
| 25 |
+
2. Facilitating methods of harassment or bullying to intimidate, abuse, or insult others;
|
| 26 |
+
3. Generating content that facilitates, promotes, or incites violence;
|
| 27 |
+
4. Generating content that facilitates, promotes, or encourages self harm;
|
| 28 |
+
5. Generating personally identifying information for distribution or other harms;
|
| 29 |
+
6. Tracking or monitoring people without their consent;
|
| 30 |
+
7. Generating content that may have unfair or adverse impacts on people, particularly impacts related to sensitive or protected characteristics; or
|
| 31 |
+
8. Generating, gathering, processing, or inferring sensitive personal or private information about individuals without obtaining all rights, authorizations, and consents required by applicable laws.
|
| 32 |
+
3. Generate and distribute content intended to misinform, misrepresent or mislead, including:
|
| 33 |
+
1. Misrepresentation of the provenance of generated content by claiming content was created by a human, or represent generated content as original works, in order to deceive;
|
| 34 |
+
2. Generation of content that impersonates an individual (living or dead) without explicit disclosure, in order to deceive;
|
| 35 |
+
3. Misleading claims of expertise or capability made particularly in sensitive areas (e.g. health, finance, government services, or legal);
|
| 36 |
+
4. Making automated decisions in domains that affect material or individual rights or well-being (e.g., finance, legal, employment, healthcare, housing, insurance, and social welfare);
|
| 37 |
+
5. Generation of defamatory content, including defamatory statements, images, or audio content; or
|
| 38 |
+
6. Engaging in the unauthorized or unlicensed practice of any profession including, but not limited to, financial, legal, medical/health, or related professional practices.
|
| 39 |
+
4. Generate sexually explicit content, including content created for the purposes of pornography or sexual gratification (e.g. sexual chatbots). Note that this does not include content created for scientific, educational, documentary, or artistic purposes.
|
LICENSE
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Gemma Terms of Use
|
| 2 |
+
Source: https://ai.google.dev/gemma/terms
|
| 3 |
+
Retrieved: 2026-09-01
|
| 4 |
+
|
| 5 |
+
The terms below apply to Gemma models listed in the Appendix at bottom of this page. For Gemma 4 terms, see the Gemma 4 license.
|
| 6 |
+
|
| 7 |
+
Last modified: April 1, 2026
|
| 8 |
+
|
| 9 |
+
By using, reproducing, modifying, distributing, performing or displaying any portion or element of Gemma, Model Derivatives including via any Hosted Service, (each as defined below) (collectively, the "Gemma Services") or otherwise accepting the terms of this Agreement, you agree to be bound by this Agreement.
|
| 10 |
+
|
| 11 |
+
SECTION 1: DEFINITIONS
|
| 12 |
+
|
| 13 |
+
1.1 Definitions
|
| 14 |
+
|
| 15 |
+
(a) "Agreement" or "Gemma Terms of Use" means these terms and conditions that govern the use, reproduction, Distribution or modification of the Gemma Services and any terms and conditions incorporated by reference.
|
| 16 |
+
|
| 17 |
+
(b) "Distribution" or "Distribute" means any transmission, publication, or other sharing of Gemma or Model Derivatives to a third party, including by providing or making Gemma or its functionality available as a hosted service via API, web access, or any other electronic or remote means ("Hosted Service").
|
| 18 |
+
|
| 19 |
+
(c) "Gemma" means the set of machine learning language models, trained model weights and parameters identified in the Appendix, regardless of the source that you obtained it from.
|
| 20 |
+
|
| 21 |
+
(d) "Google" means Google LLC.
|
| 22 |
+
|
| 23 |
+
(e) "Model Derivatives" means all (i) modifications to Gemma, (ii) works based on Gemma, or (iii) any other machine learning model which is created by transfer of patterns of the weights, parameters, operations, or Output of Gemma, to that model in order to cause that model to perform similarly to Gemma, including distillation methods that use intermediate data representations or methods based on the generation of synthetic data Outputs by Gemma for training that model. For clarity, Outputs are not deemed Model Derivatives.
|
| 24 |
+
|
| 25 |
+
(f) "Output" means the information content output of Gemma or a Model Derivative that results from operating or otherwise using Gemma or the Model Derivative, including via a Hosted Service.
|
| 26 |
+
|
| 27 |
+
1.2
|
| 28 |
+
|
| 29 |
+
As used in this Agreement, "including" means "including without limitation".
|
| 30 |
+
|
| 31 |
+
SECTION 2: ELIGIBILITY AND USAGE
|
| 32 |
+
|
| 33 |
+
2.1 Eligibility
|
| 34 |
+
|
| 35 |
+
You represent and warrant that you have the legal capacity to enter into this Agreement (including being of sufficient age of consent). If you are accessing or using any of the Gemma Services for or on behalf of a legal entity, (a) you are entering into this Agreement on behalf of yourself and that legal entity, (b) you represent and warrant that you have the authority to act on behalf of and bind that entity to this Agreement and (c) references to "you" or "your" in the remainder of this Agreement refers to both you (as an individual) and that entity.
|
| 36 |
+
|
| 37 |
+
2.2 Use
|
| 38 |
+
|
| 39 |
+
You may use, reproduce, modify, Distribute, perform or display any of the Gemma Services only in accordance with the terms of this Agreement, and must not violate (or encourage or permit anyone else to violate) any term of this Agreement.
|
| 40 |
+
|
| 41 |
+
SECTION 3: DISTRIBUTION AND RESTRICTIONS
|
| 42 |
+
|
| 43 |
+
3.1 Distribution and Redistribution
|
| 44 |
+
|
| 45 |
+
You may reproduce or Distribute copies of Gemma or Model Derivatives if you meet all of the following conditions:
|
| 46 |
+
|
| 47 |
+
1. You must include the use restrictions referenced in Section 3.2 as an enforceable provision in any agreement (e.g., license agreement, terms of use, etc.) governing the use and/or distribution of Gemma or Model Derivatives and you must provide notice to subsequent users you Distribute to that Gemma or Model Derivatives are subject to the use restrictions in Section 3.2.
|
| 48 |
+
2. You must provide all third party recipients of Gemma or Model Derivatives a copy of this Agreement.
|
| 49 |
+
3. You must cause any modified files to carry prominent notices stating that you modified the files.
|
| 50 |
+
4. All Distributions (other than through a Hosted Service) must be accompanied by a "Notice" text file that contains the following notice: "Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms".
|
| 51 |
+
|
| 52 |
+
You may add your own intellectual property statement to your modifications and, except as set forth in this Section, may provide additional or different terms and conditions for use, reproduction, or Distribution of your modifications, or for any such Model Derivatives as a whole, provided your use, reproduction, modification, Distribution, performance, and display of Gemma otherwise complies with the terms and conditions of this Agreement. Any additional or different terms and conditions you impose must not conflict with the terms of this Agreement.
|
| 53 |
+
|
| 54 |
+
3.2 Use Restrictions
|
| 55 |
+
|
| 56 |
+
You must not use any of the Gemma Services:
|
| 57 |
+
|
| 58 |
+
1. for the restricted uses set forth in the Gemma Prohibited Use Policy at ai.google.dev/gemma/prohibited_use_policy ("Prohibited Use Policy"), which is hereby incorporated by reference into this Agreement; or
|
| 59 |
+
2. in violation of applicable laws and regulations.
|
| 60 |
+
|
| 61 |
+
To the maximum extent permitted by law, Google reserves the right to restrict (remotely or otherwise) usage of any of the Gemma Services that Google reasonably believes are in violation of this Agreement.
|
| 62 |
+
|
| 63 |
+
3.3 Generated Output
|
| 64 |
+
|
| 65 |
+
Google claims no rights in Outputs you generate using Gemma. You and your users are solely responsible for Outputs and their subsequent uses.
|
| 66 |
+
|
| 67 |
+
SECTION 4: ADDITIONAL PROVISIONS
|
| 68 |
+
|
| 69 |
+
4.1 Updates
|
| 70 |
+
|
| 71 |
+
Google may update Gemma from time to time.
|
| 72 |
+
|
| 73 |
+
4.2 Trademarks
|
| 74 |
+
|
| 75 |
+
Nothing in this Agreement grants you any rights to use Google's trademarks, trade names, logos or to otherwise suggest endorsement or misrepresent the relationship between you and Google. Google reserves any rights not expressly granted herein.
|
| 76 |
+
|
| 77 |
+
4.3 DISCLAIMER OF WARRANTY
|
| 78 |
+
|
| 79 |
+
UNLESS REQUIRED BY APPLICABLE LAW, THE GEMMA SERVICES, AND OUTPUTS, ARE PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, EITHER EXPRESS OR IMPLIED, INCLUDING ANY WARRANTIES OR CONDITIONS OF TITLE, NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. YOU ARE SOLELY RESPONSIBLE FOR DETERMINING THE APPROPRIATENESS OF USING, REPRODUCING, MODIFYING, PERFORMING, DISPLAYING OR DISTRIBUTING ANY OF THE GEMMA SERVICES OR OUTPUTS AND ASSUME ANY AND ALL RISKS ASSOCIATED WITH YOUR USE OR DISTRIBUTION OF ANY OF THE GEMMA SERVICES OR OUTPUTS AND YOUR EXERCISE OF RIGHTS AND PERMISSIONS UNDER THIS AGREEMENT.
|
| 80 |
+
|
| 81 |
+
4.4 LIMITATION OF LIABILITY
|
| 82 |
+
|
| 83 |
+
TO THE FULLEST EXTENT PERMITTED BY APPLICABLE LAW, IN NO EVENT AND UNDER NO LEGAL THEORY, WHETHER IN TORT (INCLUDING NEGLIGENCE), PRODUCT LIABILITY, CONTRACT, OR OTHERWISE, UNLESS REQUIRED BY APPLICABLE LAW, SHALL GOOGLE OR ITS AFFILIATES BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY DIRECT, INDIRECT, SPECIAL, INCIDENTAL, EXEMPLARY, CONSEQUENTIAL, OR PUNITIVE DAMAGES, OR LOST PROFITS OF ANY KIND ARISING FROM THIS AGREEMENT OR RELATED TO, ANY OF THE GEMMA SERVICES OR OUTPUTS EVEN IF GOOGLE OR ITS AFFILIATES HAVE BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES.
|
| 84 |
+
|
| 85 |
+
4.5 Term, Termination, and Survival
|
| 86 |
+
|
| 87 |
+
The term of this Agreement will commence upon your acceptance of this Agreement (including acceptance by your use, modification, or Distribution, reproduction, performance or display of any portion or element of the Gemma Services) and will continue in full force and effect until terminated in accordance with the terms of this Agreement. Google may terminate this Agreement if you are in breach of any term of this Agreement. Upon termination of this Agreement, you must delete and cease use and Distribution of all copies of Gemma and Model Derivatives in your possession or control. Sections 1, 2.1, 3.3, 4.2 to 4.9 shall survive the termination of this Agreement.
|
| 88 |
+
|
| 89 |
+
4.6 Governing Law and Jurisdiction
|
| 90 |
+
|
| 91 |
+
This Agreement will be governed by the laws of the State of California without regard to choice of law principles. The UN Convention on Contracts for the International Sale of Goods does not apply to this Agreement. The state and federal courts of Santa Clara County, California shall have exclusive jurisdiction of any dispute arising out of this Agreement.
|
| 92 |
+
|
| 93 |
+
4.7 Severability
|
| 94 |
+
|
| 95 |
+
If any provision of this Agreement is held to be invalid, illegal or unenforceable, the remaining provisions shall be unaffected thereby and remain valid as if such provision had not been set forth herein.
|
| 96 |
+
|
| 97 |
+
4.8 Entire Agreement
|
| 98 |
+
|
| 99 |
+
This Agreement states all the terms agreed between the parties and supersedes all other agreements between the parties as of the date of acceptance relating to its subject matter.
|
| 100 |
+
|
| 101 |
+
4.9 No Waiver
|
| 102 |
+
|
| 103 |
+
Google will not be treated as having waived any rights by not exercising (or delaying the exercise of) any rights under this Agreement.
|
| 104 |
+
|
| 105 |
+
APPENDIX
|
| 106 |
+
|
| 107 |
+
- Gemma 1
|
| 108 |
+
- Gemma 1.1
|
| 109 |
+
- Gemma 2
|
| 110 |
+
- Gemma 3
|
| 111 |
+
- Gemma 3n
|
| 112 |
+
- FunctionGemma
|
| 113 |
+
- EmbeddingGemma
|
| 114 |
+
- PaliGemma
|
| 115 |
+
- PaliGemma 2
|
| 116 |
+
- ShieldGemma
|
| 117 |
+
- ShieldGemma 2
|
| 118 |
+
- CodeGemma
|
| 119 |
+
- CodeGemma 1.1
|
| 120 |
+
- Gemma 2 JPN
|
| 121 |
+
- DataGemma RIG
|
| 122 |
+
- DataGemma RAG
|
| 123 |
+
- RecurrentGemma
|
| 124 |
+
- Gemma Scope
|
| 125 |
+
- Gemma-APS
|
| 126 |
+
- T5Gemma
|
| 127 |
+
- VaultGemma
|
| 128 |
+
- FunctionGemma
|
| 129 |
+
- T5Gemma 2
|
| 130 |
+
- TranslateGemma
|
MODIFICATIONS.md
ADDED
|
@@ -0,0 +1,49 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Modifications from the upstream model
|
| 2 |
+
|
| 3 |
+
`embeddinggemma-300m-memory-ft-v2` is not an unmodified copy of
|
| 4 |
+
`google/embeddinggemma-300m`. It is a Daecore-modified derivative and is not
|
| 5 |
+
endorsed by Google.
|
| 6 |
+
|
| 7 |
+
What changed:
|
| 8 |
+
|
| 9 |
+
- **Fine-tune.** The base weights were trained further for semantic retrieval
|
| 10 |
+
over Daecore's governed memory corpus, using query and document role
|
| 11 |
+
prefixes, graded positive supervision, hard negatives, Matryoshka
|
| 12 |
+
dimensions, and attention LoRA (rank 16, alpha 32). The adapter was merged
|
| 13 |
+
into the model weights. Relevance grades were frontier-model judgments under
|
| 14 |
+
a frozen protocol, not human annotations.
|
| 15 |
+
- **Selected successor.** A fresh upstream fit combined cleaned private
|
| 16 |
+
supervision with annotated HotpotQA, MultiDoc2Dial and FinQA evidence,
|
| 17 |
+
upstream retention (weight 2), and a within-question grade-3 preference
|
| 18 |
+
(weight 0.25). The selected merged weights are from update 1,130. The model
|
| 19 |
+
card records their actual data exposure and measured tradeoffs; the prior
|
| 20 |
+
model's counts and qualification do not describe these weights.
|
| 21 |
+
- **Export.** The merged model was exported to ONNX as one graph taking
|
| 22 |
+
`input_ids` and `attention_mask` and emitting 768-dimensional embeddings in
|
| 23 |
+
FP32, with the parameters stored as external data.
|
| 24 |
+
- **Serving contract.** `serving.json` records the runtime and sequence
|
| 25 |
+
geometry of the original export qualification: 128 query tokens, 1,024 passage
|
| 26 |
+
tokens and a CUDA batch of 32. Production batching adapts to available memory
|
| 27 |
+
within the separately qualified CPU, CUDA and Vulkan limits. Publication
|
| 28 |
+
and source activation remain separate from a successful export.
|
| 29 |
+
|
| 30 |
+
Files modified or generated by Daecore:
|
| 31 |
+
|
| 32 |
+
- `model.onnx`: the generated ONNX graph of the fine-tuned model;
|
| 33 |
+
- `model.onnx.data`: the fine-tuned parameters used by that graph;
|
| 34 |
+
- `config.json`: the export configuration of the modified model;
|
| 35 |
+
- `serving.json`: the Daecore runtime and sequence-geometry contract.
|
| 36 |
+
- `vulkan-derivation.json`: exact source and derived graph identities and the
|
| 37 |
+
bounded mask rewrite, with learned weights unchanged by that rewrite.
|
| 38 |
+
|
| 39 |
+
Unchanged: `tokenizer.json`, `tokenizer_config.json`,
|
| 40 |
+
`special_tokens_map.json`, and `added_tokens.json` are the tokenizer files of
|
| 41 |
+
the selected model export and bind the exact text-processing contract.
|
| 42 |
+
|
| 43 |
+
This distribution is subject to the Gemma Terms of Use (`LICENSE`) and the
|
| 44 |
+
Gemma Prohibited Use Policy (`GEMMA_PROHIBITED_USE_POLICY.txt`).
|
| 45 |
+
|
| 46 |
+
- **Portable shape operations.** Inferred-dimension reshape operators use
|
| 47 |
+
`allowzero=0`. A bounded integer mask absolute-value operation runs in exact
|
| 48 |
+
FP32 and casts back to its original type, allowing native Vulkan execution.
|
| 49 |
+
CPU, CUDA and Vulkan use the same derived graph.
|
NOTICE
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Daecore memory retriever: embeddinggemma-300m-memory-ft-v2
|
| 2 |
+
|
| 3 |
+
Gemma is provided under and subject to the Gemma Terms of Use found at ai.google.dev/gemma/terms
|
| 4 |
+
|
| 5 |
+
This model is a Daecore-modified derivative of google/embeddinggemma-300m
|
| 6 |
+
(Google DeepMind). It is distributed under the Gemma Terms of Use, which
|
| 7 |
+
accompany this distribution as LICENSE, together with the incorporated Gemma
|
| 8 |
+
Prohibited Use Policy (GEMMA_PROHIBITED_USE_POLICY.txt) and the modification
|
| 9 |
+
record (MODIFICATIONS.md). Use of this model and of any further derivative is
|
| 10 |
+
subject to the use restrictions in Section 3.2 of the Gemma Terms of Use.
|
| 11 |
+
Google does not endorse this derivative.
|
README.md
ADDED
|
@@ -0,0 +1,117 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: gemma
|
| 3 |
+
base_model: google/embeddinggemma-300m
|
| 4 |
+
pipeline_tag: sentence-similarity
|
| 5 |
+
language:
|
| 6 |
+
- en
|
| 7 |
+
tags:
|
| 8 |
+
- embeddinggemma
|
| 9 |
+
- dense-retrieval
|
| 10 |
+
- fine-tuned
|
| 11 |
+
- matryoshka
|
| 12 |
+
- onnx
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# EmbeddingGemma-300M memory retriever v2
|
| 16 |
+
|
| 17 |
+
An English embedding model for searching notes, procedures, decisions and technical documentation. It maps each query or passage to a normalized 768-dimensional vector and runs locally through ONNX Runtime. Fine-tuning starts from [Google's EmbeddingGemma-300M](https://huggingface.co/google/embeddinggemma-300m).
|
| 18 |
+
|
| 19 |
+
Version 2 succeeds the [first Daecore fine-tune](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v1) with new weights. Embed queries and passages with the same revision; re-embed stored passages when switching to v2.
|
| 20 |
+
|
| 21 |
+
## Quick start
|
| 22 |
+
|
| 23 |
+
Install `onnxruntime==1.24.4`, `tokenizers` and `numpy`. Keep `model.onnx`, `model.onnx.data` and the tokenizer from the same package together. This CPU example assumes that package is in `downloaded-model`:
|
| 24 |
+
|
| 25 |
+
```python
|
| 26 |
+
from pathlib import Path
|
| 27 |
+
import numpy as np
|
| 28 |
+
import onnxruntime as ort
|
| 29 |
+
from tokenizers import Tokenizer
|
| 30 |
+
|
| 31 |
+
root = Path("downloaded-model")
|
| 32 |
+
tokenizer = Tokenizer.from_file(str(root / "tokenizer.json"))
|
| 33 |
+
session = ort.InferenceSession(
|
| 34 |
+
str(root / "model.onnx"), providers=["CPUExecutionProvider"]
|
| 35 |
+
)
|
| 36 |
+
|
| 37 |
+
def embed(text, *, query=False):
|
| 38 |
+
prefix = "task: search result | query: " if query else "title: none | text: "
|
| 39 |
+
tokenizer.enable_truncation(max_length=128 if query else 1024)
|
| 40 |
+
encoded = tokenizer.encode(prefix + text)
|
| 41 |
+
return session.run(["embeddings"], {
|
| 42 |
+
"input_ids": np.array([encoded.ids], dtype=np.int64),
|
| 43 |
+
"attention_mask": np.array([encoded.attention_mask], dtype=np.int64),
|
| 44 |
+
})[0][0]
|
| 45 |
+
|
| 46 |
+
query = embed("What must happen before a database migration?", query=True)
|
| 47 |
+
passage = embed("Take a verified backup before applying the migration.")
|
| 48 |
+
print(float(query @ passage))
|
| 49 |
+
```
|
| 50 |
+
|
| 51 |
+
Higher cosine similarity means a closer match. Use the query and passage prefixes shown above. Input limits include prefixes and special tokens: **128 tokens for queries and 1,024 for passages**. The qualified output is 768 dimensions; smaller Matryoshka widths need their own quality check.
|
| 52 |
+
|
| 53 |
+
## Evaluation
|
| 54 |
+
|
| 55 |
+
Five public retrieval datasets show how much general retrieval quality survives task adaptation. Scores are **dense-only nDCG@10** over full corpora, with matched preprocessing and official relevance judgments.
|
| 56 |
+
|
| 57 |
+
| Dataset | Queries | Upstream Gemma | Daecore fine-tune |
|
| 58 |
+
|---|---:|---:|---:|
|
| 59 |
+
| SciFact | 300 | 0.7876 | 0.7783 |
|
| 60 |
+
| FiQA | 648 | 0.4741 | 0.4468 |
|
| 61 |
+
| NFCorpus | 323 | 0.3932 | 0.3890 |
|
| 62 |
+
| SciDocs | 1,000 | 0.1945 | 0.1852 |
|
| 63 |
+
| ArguAna | 1,406 | 0.6432 | 0.6259 |
|
| 64 |
+
|
| 65 |
+
The fine-tune is below upstream on all five datasets, most on FiQA (−0.0273). None of these datasets supplied training examples, but they informed development, so they are not untouched tests.
|
| 66 |
+
|
| 67 |
+
The public companion reports this revision against the previous fine-tune inside the full Daecore search pipeline: BM25, fusion and the unchanged Ettin reranker, on 970 development queries. More queries had a useful passage at every cutoff from 3 to 20, and nDCG@10 was lower. The [evaluation companion](evaluation/README.md) gives that table with its judging rules and metric code; the [retrieval pipeline overview](https://huggingface.co/Daecore) describes the components.
|
| 68 |
+
|
| 69 |
+
## Training data and objective
|
| 70 |
+
|
| 71 |
+
Private supervision teaches passage relevance for agent retrieval. A question can have several useful passages; decisive evidence is graded above useful but incomplete evidence, and similar but unhelpful passages are explicit negatives. The documents are generated organizational material plus one person's project documentation, much of it AI-written. Four models from three provider families wrote the queries without seeing a designated answer, and model judges graded relevance under a common rubric. Construction, chunking, query writing and judging follow shared processes, so the data does not stand in for other users' workspaces.
|
| 72 |
+
|
| 73 |
+
Public supervision uses existing evidence annotations from [HotpotQA](https://hotpotqa.github.io/), [MultiDoc2Dial](https://github.com/doc2dial/multidoc2dial) and [FinQA](https://github.com/czyssrs/FinQA): multi-document questions, document-grounded dialogue, and textual or table evidence for numerical questions. The model learns to retrieve that evidence, not to perform FinQA's calculations; FinQA is unrelated to the FiQA evaluation above.
|
| 74 |
+
|
| 75 |
+
The objective combines supervised retrieval, an upstream-similarity retention penalty that limits drift from upstream Gemma's similarity structure, and a small preference for grade-3 over grade-2 private positives. Known positives and related passages are masked so they do not act as ordinary in-batch negatives.
|
| 76 |
+
|
| 77 |
+
The run completed 2,260 updates, and validation selected **update 1,130**. Counts are unique questions and query–passage pairs seen by that checkpoint:
|
| 78 |
+
|
| 79 |
+
| Source | Available questions | Questions seen | Positive pairs seen |
|
| 80 |
+
|---|---:|---:|---:|
|
| 81 |
+
| Private Daecore | 2,165 | 2,165 | 73,130 |
|
| 82 |
+
| HotpotQA | 90,272 | 45,140 | 90,280 |
|
| 83 |
+
| MultiDoc2Dial | 4,931 | 2,466 | 2,508 |
|
| 84 |
+
| FinQA | 5,754 | 2,905 | 4,966 |
|
| 85 |
+
|
| 86 |
+
The checkpoint also saw 59,086 unique explicit private negative pairs; public groups used masked in-batch comparisons instead. Question and source weighting keep the largest pool from dominating through size alone.
|
| 87 |
+
|
| 88 |
+
<details>
|
| 89 |
+
<summary>Training recipe</summary>
|
| 90 |
+
|
| 91 |
+
| Setting | Value |
|
| 92 |
+
|---|---|
|
| 93 |
+
| Initialization | Upstream Gemma; fresh optimizer |
|
| 94 |
+
| Adapter | Attention LoRA, rank 16, alpha 32; merged for serving |
|
| 95 |
+
| Effective batch | 64 question groups; up to four positives and four explicit negatives per group |
|
| 96 |
+
| Source weights | Private 2/3; each public source 1/9 |
|
| 97 |
+
| Optimizer | AdamW, learning rate 5e-5, weight decay 0.01 |
|
| 98 |
+
| Schedule | 10% warmup, cosine decay over 2,260 updates |
|
| 99 |
+
| Precision | FP16 with dynamic loss scaling |
|
| 100 |
+
| Retention / graded preference weights | 2.0 / 0.25 |
|
| 101 |
+
| Embedding widths and loss weights | 768 / 512 / 256 / 128, weighted 1 / 0.25 / 0.125 / 0.0625 |
|
| 102 |
+
|
| 103 |
+
Small validation checks selected the checkpoint before full scoring. The recipe was run once; variation across training seeds is unmeasured.
|
| 104 |
+
|
| 105 |
+
</details>
|
| 106 |
+
|
| 107 |
+
## Runtime and limits
|
| 108 |
+
|
| 109 |
+
The FP32 ONNX graph includes mean pooling, the learned projection and normalization; graph and external weights total about 1.23 GB. CPU, CUDA and Vulkan passed a 74-vector check covering token limits, mixed lengths and concurrent query and passage calls, with maximum differences from the Torch reference below 5.5e-7 and unchanged rankings on the tested inputs. Bounded mask operations were rewritten for Vulkan; learned weights are unchanged.
|
| 110 |
+
|
| 111 |
+
Vulkan runs through the native ONNX Runtime WebGPU plugin, [`onnxruntime-ep-webgpu==0.4.0`](https://pypi.org/project/onnxruntime-ep-webgpu/), with `onnxruntime==1.24.4`. Register the plugin library, add its device to the session options and create the session without a providers list. Set `dawnBackendType` to `Vulkan`, because Windows can otherwise select Direct3D 12; the tested options also set `enableInt64=1`, `powerPreference=high-performance`, `validationMode=basic` and `storageBufferCacheMode=lazyRelease`. Checks used Windows x64 with an NVIDIA RTX 3060 Ti (8 GB); other GPUs and Linux are untested, even where Vulkan is available.
|
| 112 |
+
|
| 113 |
+
The task-specific data is mostly generated and model-judged, with no human reference panel, so it describes this development distribution rather than other users' documents. Passage relevance does not measure distinct-fact coverage or downstream agent success.
|
| 114 |
+
|
| 115 |
+
## License
|
| 116 |
+
|
| 117 |
+
[Gemma Terms of Use](https://ai.google.dev/gemma/terms) and the Gemma Prohibited Use Policy. The package includes the required terms, attribution and modification notice.
|
added_tokens.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"<image_soft_token>": 262144
|
| 3 |
+
}
|
config.json
ADDED
|
@@ -0,0 +1,61 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_sliding_window_pattern": 6,
|
| 3 |
+
"architectures": [
|
| 4 |
+
"Gemma3TextModel"
|
| 5 |
+
],
|
| 6 |
+
"attention_bias": false,
|
| 7 |
+
"attention_dropout": 0.0,
|
| 8 |
+
"attn_logit_softcapping": null,
|
| 9 |
+
"bos_token_id": 2,
|
| 10 |
+
"dtype": "float32",
|
| 11 |
+
"eos_token_id": 1,
|
| 12 |
+
"final_logit_softcapping": null,
|
| 13 |
+
"head_dim": 256,
|
| 14 |
+
"hidden_activation": "gelu_pytorch_tanh",
|
| 15 |
+
"hidden_size": 768,
|
| 16 |
+
"initializer_range": 0.02,
|
| 17 |
+
"intermediate_size": 1152,
|
| 18 |
+
"layer_types": [
|
| 19 |
+
"sliding_attention",
|
| 20 |
+
"sliding_attention",
|
| 21 |
+
"sliding_attention",
|
| 22 |
+
"sliding_attention",
|
| 23 |
+
"sliding_attention",
|
| 24 |
+
"full_attention",
|
| 25 |
+
"sliding_attention",
|
| 26 |
+
"sliding_attention",
|
| 27 |
+
"sliding_attention",
|
| 28 |
+
"sliding_attention",
|
| 29 |
+
"sliding_attention",
|
| 30 |
+
"full_attention",
|
| 31 |
+
"sliding_attention",
|
| 32 |
+
"sliding_attention",
|
| 33 |
+
"sliding_attention",
|
| 34 |
+
"sliding_attention",
|
| 35 |
+
"sliding_attention",
|
| 36 |
+
"full_attention",
|
| 37 |
+
"sliding_attention",
|
| 38 |
+
"sliding_attention",
|
| 39 |
+
"sliding_attention",
|
| 40 |
+
"sliding_attention",
|
| 41 |
+
"sliding_attention",
|
| 42 |
+
"full_attention"
|
| 43 |
+
],
|
| 44 |
+
"max_position_embeddings": 2048,
|
| 45 |
+
"model_type": "gemma3_text",
|
| 46 |
+
"num_attention_heads": 3,
|
| 47 |
+
"num_hidden_layers": 24,
|
| 48 |
+
"num_key_value_heads": 1,
|
| 49 |
+
"pad_token_id": 0,
|
| 50 |
+
"query_pre_attn_scalar": 256,
|
| 51 |
+
"rms_norm_eps": 1e-06,
|
| 52 |
+
"rope_local_base_freq": 10000.0,
|
| 53 |
+
"rope_scaling": null,
|
| 54 |
+
"rope_theta": 1000000.0,
|
| 55 |
+
"sliding_window": 512,
|
| 56 |
+
"transformers_version": "4.57.6",
|
| 57 |
+
"use_bidirectional_attention": true,
|
| 58 |
+
"use_cache": false,
|
| 59 |
+
"vocab_size": 262144,
|
| 60 |
+
"attn_implementation": null
|
| 61 |
+
}
|
evaluation/README.md
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Model evaluation records
|
| 2 |
+
|
| 3 |
+
These files let readers recompute the comparisons in the Daecore model cards without model weights. Task-specific records contain anonymized row and group IDs, labels, and scores or ranked grades; none contains query or document text, private source identifiers or workspace paths. Model identities are recorded inside each file.
|
| 4 |
+
|
| 5 |
+
## Recompute
|
| 6 |
+
|
| 7 |
+
Python 3.11 or later is enough; no packages or network access are needed. From the directory that holds the records:
|
| 8 |
+
|
| 9 |
+
```sh
|
| 10 |
+
python metrics.py --directory .
|
| 11 |
+
```
|
| 12 |
+
|
| 13 |
+
The command summarizes every record present and stops with an error if there is none. Each entry in its output equals the matching `*-summary.json`. The classifier and Ettin packages also include `figures.py`; `python figures.py --output <directory>` redraws their card chart. Gemma uses a table. Each model package contains only its own records and summaries; the source repository holds all four record types. Fractions are stored at full precision; cards round percentages to two decimals and ranking metrics to four.
|
| 14 |
+
|
| 15 |
+
The records support metric recomputation. Re-running inference would need the private text and corpus, which are not included. They come from task-specific development evaluations, not a new untouched test.
|
| 16 |
+
|
| 17 |
+
## Records
|
| 18 |
+
|
| 19 |
+
| Record | Inputs and denominator | Comparison |
|
| 20 |
+
|---|---|---|
|
| 21 |
+
| `classifier.json` | 1,300 generated passages from three held-out source families; resolved labels only, per facet | Upstream GLiClass, a trained word-feature baseline and the fine-tune |
|
| 22 |
+
| `reranker.json` | 970 fixed pools of 50 passages, all 48,500 pairs judged; 873 pools have useful evidence | Upstream Ettin and the fine-tune, plus the expected result of random ordering |
|
| 23 |
+
| `promotion.json` | 970 hybrid-search queries over 82,719 passage texts, with reviewed labels and declared exclusions; five public dense-retrieval panels | Previous and updated Gemma with unchanged Ettin; the public panels also include upstream Gemma |
|
| 24 |
+
| `serving.json` | The same 970 queries with updated-Gemma candidate pools; reference extended by 32 grades | Ettin through CUDA FP16 and Vulkan FP32, with identical weights and score mapping |
|
| 25 |
+
|
| 26 |
+
`serving-qualification.json` summarizes provider and recovery checks by receipt hash and keeps the aggregate FiQA and SciFact results for upstream and fine-tuned Ettin.
|
| 27 |
+
|
| 28 |
+
“Upstream” means no Daecore fine-tuning, not an untrained network; upstream Ettin is already a trained reranker. Interim training checkpoints are not included.
|
| 29 |
+
|
| 30 |
+
## How each comparison was run
|
| 31 |
+
|
| 32 |
+
**Classifier.** Evaluation families were excluded from training. Upstream uses the same five label definitions and 768-token input limit as the fine-tune. Its raw logits and the fine-tune's calibrated probabilities are used only to rank within each facet; their scales are not compared. All 1,300 saved predictions matched fresh CPU inference of the released graph within 5.62e-6. The baseline uses word unigram and bigram TF-IDF with one balanced logistic-regression model per facet, fitted on the same 59,886 training passages; it sees full passage text, while the neural models apply their 768-token limit.
|
| 33 |
+
|
| 34 |
+
**Ettin, fixed pools.** Candidates and grades are held fixed. Both models run in PyTorch FP16 with the same 1,153-token pair construction, Transformers 5.2.0 and Sentence Transformers 5.5.1; upstream's metadata names a newer library version, but both use the same reference runtime here. The fine-tune's saved scores were checked against fresh inference on three complete pools. This evaluates the trained models, not ONNX serving or latency. The public FiQA and SciFact results rerank fixed 50-candidate pools from upstream Gemma.
|
| 35 |
+
|
| 36 |
+
**Gemma update.** Only Gemma changes; BM25, fusion and Ettin's weights stay fixed. All 970 queries are retained, and seven reviewed grade corrections apply to both models. Within each original top-k or selected prefix, 65 unresolved query–passage abstentions are excluded without backfilling: `ranked_grades` keeps their positions as `null`, with the required `excluded` flag. Precision is pooled over retained positions; nDCG reindexes them and takes its ideal ordering at the retained depth from `reference_grade_counts`. Selected depths refer to the original score-selected prefixes.
|
| 37 |
+
|
| 38 |
+
| Metric, 970 queries | Previous Gemma | Updated Gemma |
|
| 39 |
+
|---|---:|---:|
|
| 40 |
+
| Hit@3 | 83.92% | 85.36% |
|
| 41 |
+
| Hit@5 | 85.88% | 88.45% |
|
| 42 |
+
| Hit@10 | 87.84% | 89.90% |
|
| 43 |
+
| Hit@20 | 90.10% | 91.03% |
|
| 44 |
+
| nDCG@10 | 0.6764 | 0.6611 |
|
| 45 |
+
| Selected-prefix precision | 80.41% | 80.34% |
|
| 46 |
+
|
| 47 |
+
The same record holds per-query dense nDCG@10 for the five public panels. The previous fine-tune was measured only on FiQA (0.4009) and SciFact (0.7679); missing baselines stay absent. These sets supplied no training examples but informed development, and recomputing their means is different from rerunning retrieval on the public corpora.
|
| 48 |
+
|
| 49 |
+
**Ettin providers.** Updated-Gemma candidate pools for the same 970 queries were scored through CUDA FP16 and Vulkan FP32 with unchanged weights and score mapping. The two paths' top-20 results included 32 query–passage pairs without a grade. Of these, 21 reuse grades from the hard-contrast Ettin evaluation. The other 11 received two independent GPT-6 Sol judgments plus a resolution step and were then reviewed against the full passage text by GPT-6 Astra; no person reviewed them. No existing grade changed, and the 65 abstentions remain excluded. The added grades slightly change nDCG's ideal ordering, so the promotion record keeps its original reference.
|
| 50 |
+
|
| 51 |
+
| Metric, 970 queries | CUDA FP16 | Vulkan FP32 |
|
| 52 |
+
|---|---:|---:|
|
| 53 |
+
| Hit@3 | 85.36% | 85.26% |
|
| 54 |
+
| Hit@5 | 88.45% | 88.45% |
|
| 55 |
+
| Hit@10 | 89.90% | 89.90% |
|
| 56 |
+
| Hit@20 | 91.03% | 91.03% |
|
| 57 |
+
| nDCG@10 | 0.6611 | 0.6615 |
|
| 58 |
+
| Selected-prefix precision | 80.34% | 80.28% |
|
| 59 |
+
|
| 60 |
+
Small numerical differences between the paths can reorder close scores. On 20 matched pools replayed twice, second-pass reranking took 1.56 s median and 2.11 s at the 95th percentile with Vulkan, against 1.92 s and 2.70 s with the previous package's DirectML graph. All provider measurements come from one Windows x64 machine with an NVIDIA RTX 3060 Ti (8 GB) and do not transfer to other GPUs or platforms.
|
| 61 |
+
|
| 62 |
+
## Metric definitions
|
| 63 |
+
|
| 64 |
+
Relevance grades are 0–3, and grades **2 and 3** count as useful. An “answerable” query has at least one useful labeled passage in the specified pool; the absence of a useful judgment does not prove that no answer exists.
|
| 65 |
+
|
| 66 |
+
- **Hit@k:** fraction of queries with at least one useful passage in the first k positions.
|
| 67 |
+
- **Precision@k:** useful passages among the first k. Fixed-pool tables average it over queries; the promotion and serving records pool it over retained positions. Selected-prefix precision applies the same pooling to all returned passages.
|
| 68 |
+
- **Recall@k:** useful passages retrieved divided by the query's known useful passages, averaged over queries. It counts labeled passages, not every fact an answer needs.
|
| 69 |
+
- **nDCG@k:** gain `2**grade - 1`, discount `1/log2(rank + 1)`, divided by the ideal ordering at the same cutoff. Grade 1 adds a small gain although it is not useful for Hit, precision or recall. Each cutoff has its own ideal, so values need not change monotonically with k.
|
| 70 |
+
- **Average precision (AP):** area under the stepwise precision–recall curve, with tied scores grouped at one threshold; macro AP weights the five facets equally.
|
| 71 |
+
- **Precision at 90% or 95% recall:** the best measured precision at any threshold reaching that recall, taken from the evaluation curve rather than a threshold chosen in advance.
|
| 72 |
+
|
| 73 |
+
Ties keep the original candidate order. Classifier fields marked `null` are excluded identically for every model. Conditional reranker nDCG and recall use the 873 answerable pools; all-query Hit and precision use all 970. The random-order reference is exact within each pool: with N candidates and R useful passages, expected precision is R/N and Hit@k is `1 - C(N-R, k)/C(N, k)`. Random ordering already reaches 97.30% Hit@20 on the answerable reranker pools.
|
| 74 |
+
|
| 75 |
+
## Limits
|
| 76 |
+
|
| 77 |
+
Task-specific labels are language-model judgments without a human-adjudicated reference. Much of the source material is generated. Project-document sources are narrow and often AI-written; they do not establish generalization across users. The panels were reused during development, and each released model was trained once. Several relevant passages can repeat one fact, so passage-level scores do not measure unique-fact coverage, answer completeness or downstream agent success.
|
| 78 |
+
|
| 79 |
+
## Cards
|
| 80 |
+
|
| 81 |
+
- [Five-facet classifier](https://huggingface.co/Daecore/gliclass-std-base-v3-5facet-qint8-v2)
|
| 82 |
+
- [EmbeddingGemma retriever](https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v1)
|
| 83 |
+
- [Ettin reranker](https://huggingface.co/Daecore/ettin-150m-memory-reranker-ft-v1)
|
evaluation/metrics.py
ADDED
|
@@ -0,0 +1,219 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Recompute the published model comparisons from text-free evaluation records.
|
| 2 |
+
|
| 3 |
+
Python 3.11+, standard library only. Run: python metrics.py --directory .
|
| 4 |
+
"""
|
| 5 |
+
|
| 6 |
+
from __future__ import annotations
|
| 7 |
+
|
| 8 |
+
import argparse
|
| 9 |
+
import json
|
| 10 |
+
import math
|
| 11 |
+
from pathlib import Path
|
| 12 |
+
from statistics import mean
|
| 13 |
+
|
| 14 |
+
CUTOFFS = (1, 3, 5, 10, 20)
|
| 15 |
+
FACETS = ("trap", "decision", "constraint", "mechanism", "procedure")
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
def classification_metrics(labels: list[int], scores: list[float]) -> dict[str, float]:
|
| 19 |
+
"""Threshold-grouped AP and best precision at or above each recall target."""
|
| 20 |
+
if len(labels) != len(scores) or not labels or set(labels) - {0, 1}:
|
| 21 |
+
raise ValueError("Classification labels and scores must be aligned binary rows")
|
| 22 |
+
if not all(math.isfinite(s) for s in scores) or sum(labels) == 0:
|
| 23 |
+
raise ValueError("Classification scores must be finite with positive support")
|
| 24 |
+
order = sorted(range(len(labels)), key=lambda i: -scores[i])
|
| 25 |
+
positives = sum(labels)
|
| 26 |
+
true_positives = 0
|
| 27 |
+
previous_recall = 0.0
|
| 28 |
+
ap = 0.0
|
| 29 |
+
points = []
|
| 30 |
+
for rank, index in enumerate(order, 1):
|
| 31 |
+
true_positives += labels[index]
|
| 32 |
+
if rank < len(order) and scores[order[rank]] == scores[index]:
|
| 33 |
+
continue
|
| 34 |
+
precision = true_positives / rank
|
| 35 |
+
recall = true_positives / positives
|
| 36 |
+
ap += (recall - previous_recall) * precision
|
| 37 |
+
previous_recall = recall
|
| 38 |
+
points.append((precision, recall))
|
| 39 |
+
return {
|
| 40 |
+
"average_precision": ap,
|
| 41 |
+
"precision_at_recall_90": max(p for p, r in points if r >= 0.9),
|
| 42 |
+
"precision_at_recall_95": max(p for p, r in points if r >= 0.95),
|
| 43 |
+
"prevalence": positives / len(labels),
|
| 44 |
+
}
|
| 45 |
+
|
| 46 |
+
|
| 47 |
+
def dcg(grades: list[int], k: int) -> float:
|
| 48 |
+
return sum((2**g - 1) / math.log2(i + 2) for i, g in enumerate(grades[:k]))
|
| 49 |
+
|
| 50 |
+
|
| 51 |
+
def ranking_metrics(grades: list[int], scores: list[float], k: int) -> dict[str, float]:
|
| 52 |
+
if len(grades) != len(scores) or len(grades) < k or k < 1:
|
| 53 |
+
raise ValueError("Ranking inputs must be aligned and cover the cutoff")
|
| 54 |
+
if any(type(g) is not int or g not in range(4) for g in grades):
|
| 55 |
+
raise ValueError("Ranking requires fully judged integer grades 0 through 3")
|
| 56 |
+
if not all(math.isfinite(s) for s in scores):
|
| 57 |
+
raise ValueError("Ranking scores must be finite")
|
| 58 |
+
order = sorted(range(len(scores)), key=lambda i: (-scores[i], i))
|
| 59 |
+
ranked = [grades[i] for i in order]
|
| 60 |
+
useful = sum(g >= 2 for g in ranked[:k])
|
| 61 |
+
relevant = sum(g >= 2 for g in grades)
|
| 62 |
+
ideal = dcg(sorted(grades, reverse=True), k)
|
| 63 |
+
result = {"hit": float(useful > 0), "precision": useful / k, "useful": float(useful)}
|
| 64 |
+
if relevant:
|
| 65 |
+
result["recall"] = useful / relevant
|
| 66 |
+
result["ndcg"] = dcg(ranked, k) / ideal
|
| 67 |
+
return result
|
| 68 |
+
|
| 69 |
+
|
| 70 |
+
def random_ranking_metrics(grades: list[int], k: int) -> dict[str, float]:
|
| 71 |
+
"""Exact expectation under a uniform permutation of one fixed judged pool."""
|
| 72 |
+
n = len(grades)
|
| 73 |
+
if k < 1 or k > n or any(type(g) is not int or g not in range(4) for g in grades):
|
| 74 |
+
raise ValueError("Random reference requires judged grades and a valid cutoff")
|
| 75 |
+
relevant = sum(g >= 2 for g in grades)
|
| 76 |
+
misses = math.comb(n - relevant, k) if n - relevant >= k else 0
|
| 77 |
+
result = {"hit": 1 - misses / math.comb(n, k), "precision": relevant / n, "useful": k * relevant / n}
|
| 78 |
+
if relevant:
|
| 79 |
+
expected_dcg = mean(2**g - 1 for g in grades) * sum(1 / math.log2(i + 2) for i in range(k))
|
| 80 |
+
result["recall"] = k / n
|
| 81 |
+
result["ndcg"] = expected_dcg / dcg(sorted(grades, reverse=True), k)
|
| 82 |
+
return result
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
def summarize_classifier(data: dict) -> dict:
|
| 86 |
+
rows = data["rows"]
|
| 87 |
+
result = {"rows": len(rows), "families": len({r["group"] for r in rows}), "models": {}}
|
| 88 |
+
for model in data["model_order"]:
|
| 89 |
+
by_facet = {}
|
| 90 |
+
for facet in FACETS:
|
| 91 |
+
resolved = [r for r in rows if r["labels"][facet] is not None]
|
| 92 |
+
if any(r["labels"][facet] not in (0, 1) for r in resolved):
|
| 93 |
+
raise ValueError("Invalid resolved classifier label")
|
| 94 |
+
by_facet[facet] = {"rows": len(resolved), **classification_metrics(
|
| 95 |
+
[r["labels"][facet] for r in resolved], [r["scores"][model][facet] for r in resolved]
|
| 96 |
+
)}
|
| 97 |
+
macro = {k: mean(v[k] for v in by_facet.values()) for k in ("average_precision", "precision_at_recall_90", "precision_at_recall_95")}
|
| 98 |
+
result["models"][model] = {"facets": by_facet, "macro": macro}
|
| 99 |
+
return result
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
def summarize_reranker(data: dict) -> dict:
|
| 103 |
+
rows = data["rows"]
|
| 104 |
+
answerable = [r for r in rows if any(g >= 2 for g in r["grades"])]
|
| 105 |
+
output = {"queries": len(rows), "answerable": len(answerable), "models": {}}
|
| 106 |
+
for model in ["random", *data["model_order"]]:
|
| 107 |
+
scopes = {}
|
| 108 |
+
for scope, subset in [("answerable", answerable), ("all", rows)]:
|
| 109 |
+
cuts = {}
|
| 110 |
+
for k in CUTOFFS:
|
| 111 |
+
values = [random_ranking_metrics(r["grades"], k) if model == "random" else ranking_metrics(r["grades"], r["scores"][model], k) for r in subset]
|
| 112 |
+
fields = ("hit", "precision", "useful", "recall", "ndcg") if scope == "answerable" else ("hit", "precision", "useful")
|
| 113 |
+
cuts[str(k)] = {field: mean(v[field] for v in values) for field in fields}
|
| 114 |
+
scopes[scope] = cuts
|
| 115 |
+
output["models"][model] = scopes
|
| 116 |
+
return output
|
| 117 |
+
|
| 118 |
+
|
| 119 |
+
def summarize_promotion(data: dict) -> dict:
|
| 120 |
+
"""Hybrid search with reviewed exclusions inside each original prefix.
|
| 121 |
+
|
| 122 |
+
A null is permitted only for an explicitly excluded judging abstention.
|
| 123 |
+
Cut first, remove exclusions second, and never backfill from a deeper rank.
|
| 124 |
+
"""
|
| 125 |
+
rows = data['rows']
|
| 126 |
+
if not rows or len({row['id'] for row in rows}) != len(rows):
|
| 127 |
+
raise ValueError('Promotion rows require unique nonempty query identities')
|
| 128 |
+
output = {'queries': len(rows), 'models': {}, 'public': {}}
|
| 129 |
+
for model in data['model_order']:
|
| 130 |
+
cutoffs = {}
|
| 131 |
+
for cutoff in (3, 5, 10, 20, 'selected'):
|
| 132 |
+
per_query = []
|
| 133 |
+
for row in rows:
|
| 134 |
+
counts = row['reference_grade_counts']
|
| 135 |
+
if set(counts) != {'0', '1', '2', '3'} or any(type(n) is not int or n < 0 for n in counts.values()):
|
| 136 |
+
raise ValueError('Reference grade counts must cover grades zero through three')
|
| 137 |
+
ranked = row['ranked_grades'][model]
|
| 138 |
+
excluded = row['excluded'][model]
|
| 139 |
+
depth = row['selected_depth'][model] if cutoff == 'selected' else cutoff
|
| 140 |
+
if (len(ranked) != 20 or len(excluded) != 20 or type(depth) is not int
|
| 141 |
+
or not 3 <= depth <= 20 or any(type(x) is not bool for x in excluded)):
|
| 142 |
+
raise ValueError('Promotion rows require a bounded original top twenty')
|
| 143 |
+
if any((grade is not None if drop else type(grade) is not int or grade not in range(4))
|
| 144 |
+
for grade, drop in zip(ranked, excluded, strict=True)):
|
| 145 |
+
raise ValueError('Only declared abstentions may lack grades')
|
| 146 |
+
kept = [grade for grade, drop in zip(ranked[:depth], excluded[:depth], strict=True) if not drop]
|
| 147 |
+
if not kept:
|
| 148 |
+
raise ValueError('Every scored prefix must retain judged passages')
|
| 149 |
+
ideal_grades = [g for g in (3, 2, 1, 0) for _ in range(min(counts[str(g)], len(kept)))][:len(kept)]
|
| 150 |
+
useful = sum(g >= 2 for g in kept)
|
| 151 |
+
positives = counts['2'] + counts['3']
|
| 152 |
+
ideal = dcg(ideal_grades, len(kept))
|
| 153 |
+
per_query.append({
|
| 154 |
+
'hit': float(useful > 0), 'precision': useful / len(kept),
|
| 155 |
+
'useful': useful, 'retained': len(kept), 'excluded': depth - len(kept),
|
| 156 |
+
'ndcg': dcg(kept, len(kept)) / ideal if ideal else 0.0,
|
| 157 |
+
'known_recall': useful / positives if positives else None,
|
| 158 |
+
})
|
| 159 |
+
useful = sum(row['useful'] for row in per_query)
|
| 160 |
+
retained = sum(row['retained'] for row in per_query)
|
| 161 |
+
recalls = [row['known_recall'] for row in per_query if row['known_recall'] is not None]
|
| 162 |
+
cutoffs[str(cutoff)] = {
|
| 163 |
+
'hit': mean(row['hit'] for row in per_query), 'precision': useful / retained,
|
| 164 |
+
'macro_precision': mean(row['precision'] for row in per_query),
|
| 165 |
+
'ndcg': mean(row['ndcg'] for row in per_query),
|
| 166 |
+
'known_positive_recall': mean(recalls) if recalls else None,
|
| 167 |
+
'recall_queries': len(recalls), 'useful': useful, 'retained': retained,
|
| 168 |
+
'excluded_positions': sum(row['excluded'] for row in per_query),
|
| 169 |
+
'mean_useful': useful / len(rows), 'mean_retained': retained / len(rows),
|
| 170 |
+
}
|
| 171 |
+
output['models'][model] = cutoffs
|
| 172 |
+
for dataset, panel in data['public'].items():
|
| 173 |
+
if not panel['rows'] or len({row['id'] for row in panel['rows']}) != len(panel['rows']):
|
| 174 |
+
raise ValueError('Public panel requires distinct query identities')
|
| 175 |
+
measured = panel['model_order']
|
| 176 |
+
for row in panel['rows']:
|
| 177 |
+
if set(row['ndcg@10']) != set(measured) or any(
|
| 178 |
+
isinstance(value, bool) or not isinstance(value, int | float)
|
| 179 |
+
or not math.isfinite(value) or not 0 <= value <= 1
|
| 180 |
+
for value in row['ndcg@10'].values()
|
| 181 |
+
):
|
| 182 |
+
raise ValueError('Public nDCG values must be finite measured scores')
|
| 183 |
+
output['public'][dataset] = {
|
| 184 |
+
'queries': len(panel['rows']),
|
| 185 |
+
'ndcg@10': {model: mean(row['ndcg@10'][model] for row in panel['rows']) for model in measured},
|
| 186 |
+
}
|
| 187 |
+
return output
|
| 188 |
+
|
| 189 |
+
|
| 190 |
+
def main() -> None:
|
| 191 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 192 |
+
parser.add_argument("--directory", type=Path, default=Path(__file__).resolve().parent)
|
| 193 |
+
parser.add_argument("--output", type=Path)
|
| 194 |
+
parser.add_argument("--model", choices=("classifier", "reranker", "promotion", "serving"), help="Recompute one comparison; default: every record in this package")
|
| 195 |
+
args = parser.parse_args()
|
| 196 |
+
summarizers = {"classifier": summarize_classifier, "reranker": summarize_reranker,
|
| 197 |
+
"promotion": summarize_promotion, "serving": summarize_promotion}
|
| 198 |
+
names = [args.model] if args.model else [
|
| 199 |
+
name for name in summarizers if (args.directory / f"{name}.json").is_file()
|
| 200 |
+
]
|
| 201 |
+
if not names:
|
| 202 |
+
parser.error("No evaluation records found in the selected directory")
|
| 203 |
+
result = {}
|
| 204 |
+
for name in names:
|
| 205 |
+
summarize = summarizers[name]
|
| 206 |
+
data = json.loads((args.directory / f"{name}.json").read_text(encoding="utf-8"))
|
| 207 |
+
ids = [r["id"] for r in data["rows"]]
|
| 208 |
+
if len(set(ids)) != len(ids):
|
| 209 |
+
raise ValueError(f"Repeated query/passage identity in {name}")
|
| 210 |
+
result[name] = summarize(data)
|
| 211 |
+
text = json.dumps(result, indent=2, allow_nan=False) + "\n"
|
| 212 |
+
if args.output:
|
| 213 |
+
args.output.write_text(text, encoding="utf-8")
|
| 214 |
+
else:
|
| 215 |
+
print(text, end="")
|
| 216 |
+
|
| 217 |
+
|
| 218 |
+
if __name__ == "__main__":
|
| 219 |
+
main()
|
evaluation/promotion-summary.json
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"queries":970,"models":{"production":{"3":{"hit":0.8391752577319588,"precision":0.7108890420399724,"macro_precision":0.7109965635738832,"ndcg":0.6737359241548024,"known_positive_recall":0.09151322333990437,"recall_queries":895,"useful":2063,"retained":2902,"excluded_positions":8,"mean_useful":2.12680412371134,"mean_retained":2.9917525773195877},"5":{"hit":0.8587628865979381,"precision":0.6968194960760017,"macro_precision":0.6970103092783505,"ndcg":0.6720934096988599,"known_positive_recall":0.14456571515697228,"recall_queries":895,"useful":3374,"retained":4842,"excluded_positions":8,"mean_useful":3.4783505154639176,"mean_retained":4.991752577319588},"10":{"hit":0.8783505154639175,"precision":0.6640181611804767,"macro_precision":0.6641809851088202,"ndcg":0.6764288375969912,"known_positive_recall":0.2647168884118442,"recall_queries":895,"useful":6435,"retained":9691,"excluded_positions":9,"mean_useful":6.634020618556701,"mean_retained":9.990721649484536},"20":{"hit":0.9010309278350516,"precision":0.6041172221648953,"macro_precision":0.604256841821554,"ndcg":0.6943443750382963,"known_positive_recall":0.4574631459961227,"recall_queries":895,"useful":11709,"retained":19382,"excluded_positions":18,"mean_useful":12.071134020618556,"mean_retained":19.981443298969072},"selected":{"hit":0.8525773195876288,"precision":0.8040927303949628,"macro_precision":0.7015133181852246,"ndcg":0.67618405088445,"known_positive_recall":0.19797868631055085,"recall_queries":895,"useful":5619,"retained":6988,"excluded_positions":8,"mean_useful":5.792783505154639,"mean_retained":7.204123711340206}},"candidate":{"3":{"hit":0.8536082474226804,"precision":0.7157640565712314,"macro_precision":0.7166666666666667,"ndcg":0.6576832474683237,"known_positive_recall":0.09526530468586906,"recall_queries":895,"useful":2075,"retained":2899,"excluded_positions":11,"mean_useful":2.1391752577319587,"mean_retained":2.988659793814433},"5":{"hit":0.8845360824742268,"precision":0.7028311634635255,"macro_precision":0.7031786941580757,"ndcg":0.6554718901879515,"known_positive_recall":0.15391852985862914,"recall_queries":895,"useful":3401,"retained":4839,"excluded_positions":11,"mean_useful":3.5061855670103093,"mean_retained":4.988659793814433},"10":{"hit":0.8989690721649485,"precision":0.67180070291503,"macro_precision":0.6722631320569465,"ndcg":0.6610872123112499,"known_positive_recall":0.28045885967323153,"recall_queries":895,"useful":6499,"retained":9674,"excluded_positions":26,"mean_useful":6.7,"mean_retained":9.97319587628866},"20":{"hit":0.9103092783505154,"precision":0.61794500723589,"macro_precision":0.6184956360634603,"ndcg":0.684223528795109,"known_positive_recall":0.4887591886594314,"recall_queries":895,"useful":11956,"retained":19348,"excluded_positions":52,"mean_useful":12.32577319587629,"mean_retained":19.94639175257732},"selected":{"hit":0.8618556701030928,"precision":0.8033770583310076,"macro_precision":0.7040325511660611,"ndcg":0.6599128936120324,"known_positive_recall":0.20805438995183476,"recall_queries":895,"useful":5757,"retained":7166,"excluded_positions":17,"mean_useful":5.935051546391753,"mean_retained":7.387628865979382}}},"public":{"scifact":{"queries":300,"ndcg@10":{"upstream":0.7875641458078616,"production":0.7678535660256963,"candidate":0.7782845713792622}},"fiqa":{"queries":648,"ndcg@10":{"upstream":0.474145026242454,"production":0.4008741437969072,"candidate":0.446838124778398}},"nfcorpus":{"queries":323,"ndcg@10":{"upstream":0.3932480325545885,"candidate":0.3889646280478944}},"scidocs":{"queries":1000,"ndcg@10":{"upstream":0.19447454865159183,"candidate":0.18521860884727676}},"arguana":{"queries":1406,"ndcg@10":{"upstream":0.6431541443325537,"candidate":0.6258634411077135}}}}
|
evaluation/promotion.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
evaluation/serving-qualification.json
ADDED
|
@@ -0,0 +1,134 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema": "daecore.serving-preparation-summary.v1",
|
| 3 |
+
"date": "2026-09-28",
|
| 4 |
+
"status": "local-model-serving-checks-complete-release-transition-pending",
|
| 5 |
+
"hardware": "Windows x64, NVIDIA RTX 3060 Ti 8 GiB",
|
| 6 |
+
"unmeasured_targets": [
|
| 7 |
+
"AMD",
|
| 8 |
+
"Intel",
|
| 9 |
+
"Linux x64"
|
| 10 |
+
],
|
| 11 |
+
"runtimes": {
|
| 12 |
+
"onnxruntime": "1.24.4",
|
| 13 |
+
"vulkan_plugin": "0.4.0",
|
| 14 |
+
"backend": "Vulkan",
|
| 15 |
+
"storage_buffer_cache": "lazyRelease"
|
| 16 |
+
},
|
| 17 |
+
"models": {
|
| 18 |
+
"gemma": {
|
| 19 |
+
"vectors": 74,
|
| 20 |
+
"providers": {
|
| 21 |
+
"cpu": {
|
| 22 |
+
"receipt_sha256": "31851791104fa95d6b7b25d35f950b25e239a3d6f8fa458addb2a44917c25cb4",
|
| 23 |
+
"maximum_vector_delta": 3.0174851417541504e-07,
|
| 24 |
+
"concurrent_delta": 0.0
|
| 25 |
+
},
|
| 26 |
+
"cuda": {
|
| 27 |
+
"receipt_sha256": "6977d0aa0ebf9c39ea0fd24078c47367d1d50e81184653225fcd10758961903e",
|
| 28 |
+
"maximum_vector_delta": 2.644956111907959e-07,
|
| 29 |
+
"concurrent_delta": 2.2351741790771484e-08
|
| 30 |
+
},
|
| 31 |
+
"vulkan": {
|
| 32 |
+
"receipt_sha256": "567fb89ea195e8b188203d2334f3eda82994ae8c7fc97f2126fac3e6f4766079",
|
| 33 |
+
"maximum_vector_delta": 5.438923835754395e-07,
|
| 34 |
+
"concurrent_delta": 0.0
|
| 35 |
+
}
|
| 36 |
+
},
|
| 37 |
+
"admission": {
|
| 38 |
+
"cases": 3841,
|
| 39 |
+
"semantic_fallbacks": 679,
|
| 40 |
+
"new_useful_exclusions": 0,
|
| 41 |
+
"additional_noise_retained": 2,
|
| 42 |
+
"receipt_sha256": "71550f606c95c1e1192dd5939488b280d5f2f2f091f99413f6b63f4408a5bc2c"
|
| 43 |
+
}
|
| 44 |
+
},
|
| 45 |
+
"classifier": {
|
| 46 |
+
"cpu": {
|
| 47 |
+
"rows": 1300,
|
| 48 |
+
"maximum_posterior_delta": 5.612167303103988e-06,
|
| 49 |
+
"changed_threshold_labels": {
|
| 50 |
+
"contract": 0,
|
| 51 |
+
"recall_leaning": 0
|
| 52 |
+
},
|
| 53 |
+
"receipt_sha256": "ea545fb0a30ff25bf3da178b60593e934e6025e789c2db67290ff5360e8669a4"
|
| 54 |
+
},
|
| 55 |
+
"cuda": {
|
| 56 |
+
"rows": 1300,
|
| 57 |
+
"maximum_posterior_delta": 5.612167303103988e-06,
|
| 58 |
+
"changed_threshold_labels": {
|
| 59 |
+
"contract": 0,
|
| 60 |
+
"recall_leaning": 0
|
| 61 |
+
},
|
| 62 |
+
"receipt_sha256": "5616891e9dfdc3b6615951901cb24aa31189d6a81ed84b0080f461b85c045065"
|
| 63 |
+
},
|
| 64 |
+
"vulkan": {
|
| 65 |
+
"rows": 1300,
|
| 66 |
+
"maximum_posterior_delta": 5.612167303103988e-06,
|
| 67 |
+
"changed_threshold_labels": {
|
| 68 |
+
"contract": 0,
|
| 69 |
+
"recall_leaning": 0
|
| 70 |
+
},
|
| 71 |
+
"receipt_sha256": "c7ec971e020d9771adf0756316f468629cc05e300d7edf9abb52612901adf6b0"
|
| 72 |
+
}
|
| 73 |
+
},
|
| 74 |
+
"ettin": {
|
| 75 |
+
"queries": 970,
|
| 76 |
+
"unjudged_top20": 0,
|
| 77 |
+
"quality_evidence": "serving.json",
|
| 78 |
+
"quality_receipt_sha256": "805a1b0acd7baaf248e182908bc23cd52171e40aa8f1281799a20d15b58c6905",
|
| 79 |
+
"full_replay_seconds_p50_p95": [
|
| 80 |
+
1.7189999999827705,
|
| 81 |
+
2.610000000044238
|
| 82 |
+
],
|
| 83 |
+
"matched_20_pool_seconds_p50_p95": {
|
| 84 |
+
"vulkan": [
|
| 85 |
+
1.5565476999909151,
|
| 86 |
+
2.1063475799834124
|
| 87 |
+
],
|
| 88 |
+
"predecessor_directml": [
|
| 89 |
+
1.9197023500164505,
|
| 90 |
+
2.701977299965802
|
| 91 |
+
]
|
| 92 |
+
},
|
| 93 |
+
"pressure": "One preflight refusal after 322 queries with lazy release; resumed all remaining rows with zero further retries. An earlier default-cache run stopped after 424. No native OOM observed.",
|
| 94 |
+
"precision": "FP32 Vulkan versus FP16 CUDA; rankings not bit-exact",
|
| 95 |
+
"selector": "Unchanged CUDA mapping transferred for measurement; new package/provider binding pending"
|
| 96 |
+
}
|
| 97 |
+
},
|
| 98 |
+
"operator_runtime_changed": false,
|
| 99 |
+
"published": false,
|
| 100 |
+
"remaining": [
|
| 101 |
+
"immutable-publication-revisions",
|
| 102 |
+
"profile-and-calibration-release-bindings",
|
| 103 |
+
"isolated-package-update-and-index-rebuild-recovery",
|
| 104 |
+
"DirectML-current-route-retirement"
|
| 105 |
+
],
|
| 106 |
+
"worker_recovery": {
|
| 107 |
+
"receipt_sha256": "46205d39521bde40d1228c837e57f88e1a941d42f6846f4a244c48b513347cc6",
|
| 108 |
+
"all_three_consumer_outputs_identical_after_owned_worker_crash": true,
|
| 109 |
+
"ettin_maximum_envelope": {
|
| 110 |
+
"tokens": 1153,
|
| 111 |
+
"max_abs_logit_delta_vs_cpu": 5.91278076171875e-05
|
| 112 |
+
}
|
| 113 |
+
},
|
| 114 |
+
"retained_reranker_public": {
|
| 115 |
+
"scope": "Retained matched public pools retrieved by upstream Gemma; upstream and production Ettin PyTorch FP16; not a new Vulkan public replay",
|
| 116 |
+
"source_sha256": "0814db560081af8376e4b1f3e3d2bcd5c4d9b8b208f7fb8b7172f113e8123ef3",
|
| 117 |
+
"panels": {
|
| 118 |
+
"fiqa": {
|
| 119 |
+
"queries": 648,
|
| 120 |
+
"ndcg@10": {
|
| 121 |
+
"upstream": 0.48603492061219766,
|
| 122 |
+
"finetuned": 0.452680922415071
|
| 123 |
+
}
|
| 124 |
+
},
|
| 125 |
+
"scifact": {
|
| 126 |
+
"queries": 300,
|
| 127 |
+
"ndcg@10": {
|
| 128 |
+
"upstream": 0.7487436478294831,
|
| 129 |
+
"finetuned": 0.7543833016238701
|
| 130 |
+
}
|
| 131 |
+
}
|
| 132 |
+
}
|
| 133 |
+
}
|
| 134 |
+
}
|
model.onnx
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:04eb85c8ab0991122030856f246f4f9cb922a7c61c3045b8978a32df5d1cb4d3
|
| 3 |
+
size 4392411
|
model.onnx.data
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:56af931663f87a32abce8dd213bc87e595c1aafb12668b60408bf005cace2247
|
| 3 |
+
size 1230372864
|
publication-manifest.json
ADDED
|
@@ -0,0 +1,114 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema": "daecore.retrieval-publication-manifest",
|
| 3 |
+
"tool_sha256": "5ac2ef9465fe9a3454bfd2e61c7aa651b08fbdfef9473fb5b3455585ba560a7b",
|
| 4 |
+
"model_id": "embeddinggemma-300m-memory-ft-v2",
|
| 5 |
+
"public_repo": "Daecore/embeddinggemma-300m-memory-ft-v2",
|
| 6 |
+
"role": "dense-retriever",
|
| 7 |
+
"base": {
|
| 8 |
+
"model_id": "google/embeddinggemma-300m",
|
| 9 |
+
"revision": "kaggle:google/embeddinggemma/transformers/embeddinggemma-300m/1",
|
| 10 |
+
"tree_sha256": "24b175205b9fc9a4ba6480a54ee8ea8c26178d52ef33752d6f82d53921e96b4d"
|
| 11 |
+
},
|
| 12 |
+
"export_receipt_sha256": "e24e1692f7e03c9af8040bb3f6e310a061d305f485861007f1f37d36808b20d8",
|
| 13 |
+
"serving_sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
|
| 14 |
+
"source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
|
| 15 |
+
"staged_at": "2026-09-29T03:04:16+00:00",
|
| 16 |
+
"files": {
|
| 17 |
+
"added_tokens.json": {
|
| 18 |
+
"sha256": "50b2f405ba56a26d4913fd772089992252d7f942123cc0a034d96424221ba946",
|
| 19 |
+
"size": 35,
|
| 20 |
+
"binding": "export receipt (file hash)"
|
| 21 |
+
},
|
| 22 |
+
"config.json": {
|
| 23 |
+
"sha256": "ca1f955552087e2025d4cd2781efc3683fdd9d7f2b1b93cd67bb04eb95adfaf0",
|
| 24 |
+
"size": 1515,
|
| 25 |
+
"binding": "export receipt (file hash)"
|
| 26 |
+
},
|
| 27 |
+
"model.onnx": {
|
| 28 |
+
"sha256": "04eb85c8ab0991122030856f246f4f9cb922a7c61c3045b8978a32df5d1cb4d3",
|
| 29 |
+
"size": 4392411,
|
| 30 |
+
"binding": "export receipt (file hash)"
|
| 31 |
+
},
|
| 32 |
+
"model.onnx.data": {
|
| 33 |
+
"sha256": "56af931663f87a32abce8dd213bc87e595c1aafb12668b60408bf005cace2247",
|
| 34 |
+
"size": 1230372864,
|
| 35 |
+
"binding": "export receipt (file hash)"
|
| 36 |
+
},
|
| 37 |
+
"serving.json": {
|
| 38 |
+
"sha256": "17cf2e92caa17ad8fdd3ff37b24afe148c8bea2ecd4724e084565387ae7cd668",
|
| 39 |
+
"size": 804,
|
| 40 |
+
"binding": "export receipt (file hash)"
|
| 41 |
+
},
|
| 42 |
+
"special_tokens_map.json": {
|
| 43 |
+
"sha256": "2f7b0adf4fb469770bb1490e3e35df87b1dc578246c5e7e6fc76ecf33213a397",
|
| 44 |
+
"size": 662,
|
| 45 |
+
"binding": "export receipt (file hash)"
|
| 46 |
+
},
|
| 47 |
+
"tokenizer.json": {
|
| 48 |
+
"sha256": "216e2a79606fe879c9f17c529c71cd241338407fd5646b595ffd3c4b9ea1d503",
|
| 49 |
+
"size": 33385262,
|
| 50 |
+
"binding": "export receipt (file hash)"
|
| 51 |
+
},
|
| 52 |
+
"tokenizer_config.json": {
|
| 53 |
+
"sha256": "5cf4fdd1d8d40f8107f3fb5e3e92449c039c835d9c32e578443347a2183bc338",
|
| 54 |
+
"size": 861682,
|
| 55 |
+
"binding": "export receipt (file hash)"
|
| 56 |
+
},
|
| 57 |
+
"vulkan-derivation.json": {
|
| 58 |
+
"sha256": "3a8bd536fa3c6bf70bea7c70d3b452e3911ca1f6032ff5a1aa2bc0c1d9056fac",
|
| 59 |
+
"size": 767,
|
| 60 |
+
"binding": "export receipt (file hash)"
|
| 61 |
+
},
|
| 62 |
+
"LICENSE": {
|
| 63 |
+
"sha256": "1035e3b717a5dc47d13f36a37c1f4b7aed1dc832399c95e606a874f12551732f",
|
| 64 |
+
"size": 8916,
|
| 65 |
+
"binding": "packaging record"
|
| 66 |
+
},
|
| 67 |
+
"GEMMA_PROHIBITED_USE_POLICY.txt": {
|
| 68 |
+
"sha256": "5a6e86f6268a85896fe073e7ea64246f549778a440ebaa1358fccd0e49c18745",
|
| 69 |
+
"size": 3887,
|
| 70 |
+
"binding": "packaging record"
|
| 71 |
+
},
|
| 72 |
+
"NOTICE": {
|
| 73 |
+
"sha256": "1109569113111d00c0db230bb887f8fa4ee15dc23415f12878412b89a99467fa",
|
| 74 |
+
"size": 652,
|
| 75 |
+
"binding": "packaging record"
|
| 76 |
+
},
|
| 77 |
+
"MODIFICATIONS.md": {
|
| 78 |
+
"sha256": "bb0edbda8a9f91558e35c6c266d8c7066cd1dd1e1a902979b57706588fe6a377",
|
| 79 |
+
"size": 2739,
|
| 80 |
+
"binding": "packaging record"
|
| 81 |
+
},
|
| 82 |
+
"evaluation/README.md": {
|
| 83 |
+
"sha256": "3215e6b0e302fdecd86ef1457e46e7d033fa18fb1ebcbd0dcb47c2b8abe7c174",
|
| 84 |
+
"size": 8825,
|
| 85 |
+
"binding": "packaging record"
|
| 86 |
+
},
|
| 87 |
+
"evaluation/metrics.py": {
|
| 88 |
+
"sha256": "7c25f29a03e0f61d5f0e59f7781269489f80512ae6862920f91a2d333c058bc4",
|
| 89 |
+
"size": 11135,
|
| 90 |
+
"binding": "packaging record"
|
| 91 |
+
},
|
| 92 |
+
"evaluation/promotion.json": {
|
| 93 |
+
"sha256": "dca79b36c99e2f90c05bff12227634fe175f1370a2ac83368da084ccf7aeffc3",
|
| 94 |
+
"size": 816500,
|
| 95 |
+
"binding": "packaging record"
|
| 96 |
+
},
|
| 97 |
+
"evaluation/promotion-summary.json": {
|
| 98 |
+
"sha256": "3dcd954b778e8955f90c7e160790227602b701eea0897d539ecc1f8ab265567a",
|
| 99 |
+
"size": 3724,
|
| 100 |
+
"binding": "packaging record"
|
| 101 |
+
},
|
| 102 |
+
"evaluation/serving-qualification.json": {
|
| 103 |
+
"sha256": "bf7deda3f25306d15454798ad4c8e785c25fba970fc774efbe557225fc5d42e0",
|
| 104 |
+
"size": 4667,
|
| 105 |
+
"binding": "packaging record"
|
| 106 |
+
},
|
| 107 |
+
"README.md": {
|
| 108 |
+
"sha256": "adf7d4028bce7fa2a08764062dc748b3990981a0442b5a859dc1ba177afcc50b",
|
| 109 |
+
"size": 7829,
|
| 110 |
+
"binding": "model card with upload-relative links",
|
| 111 |
+
"source_sha256": "0849c5c0d83ab09f6c08b0f0cec3150a779c57ea1fde1d010496d5a4f9809b2d"
|
| 112 |
+
}
|
| 113 |
+
}
|
| 114 |
+
}
|
serving.json
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"external_data": "model.onnx.data",
|
| 3 |
+
"geometry": {
|
| 4 |
+
"dynamic_batch": true,
|
| 5 |
+
"dynamic_sequence": true,
|
| 6 |
+
"local_attention_window": 257,
|
| 7 |
+
"max_batch": 32,
|
| 8 |
+
"max_sequence": 1024,
|
| 9 |
+
"passage_max_sequence": 1024,
|
| 10 |
+
"query_max_sequence": 128
|
| 11 |
+
},
|
| 12 |
+
"graph": "model.onnx",
|
| 13 |
+
"inputs": [
|
| 14 |
+
"input_ids",
|
| 15 |
+
"attention_mask"
|
| 16 |
+
],
|
| 17 |
+
"output": "embeddings",
|
| 18 |
+
"precision": "float32",
|
| 19 |
+
"provider_chain": [
|
| 20 |
+
"CUDAExecutionProvider",
|
| 21 |
+
"CPUExecutionProvider"
|
| 22 |
+
],
|
| 23 |
+
"role": "dense-retriever",
|
| 24 |
+
"schema": "daecore.retrieval-onnx-serving.v1",
|
| 25 |
+
"sha256": "8c088b1635875d5406e148782b303cfc5a961138600783fb4842b413b98a98c3",
|
| 26 |
+
"source_model_tree_sha256": "30a7b9f1c42b1fcee29a6729742685d5b744e8fa87d13cefab2fd2d49ea9150e",
|
| 27 |
+
"torch_required_at_runtime": false
|
| 28 |
+
}
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"boi_token": "<start_of_image>",
|
| 3 |
+
"bos_token": {
|
| 4 |
+
"content": "<bos>",
|
| 5 |
+
"lstrip": false,
|
| 6 |
+
"normalized": false,
|
| 7 |
+
"rstrip": false,
|
| 8 |
+
"single_word": false
|
| 9 |
+
},
|
| 10 |
+
"eoi_token": "<end_of_image>",
|
| 11 |
+
"eos_token": {
|
| 12 |
+
"content": "<eos>",
|
| 13 |
+
"lstrip": false,
|
| 14 |
+
"normalized": false,
|
| 15 |
+
"rstrip": false,
|
| 16 |
+
"single_word": false
|
| 17 |
+
},
|
| 18 |
+
"image_token": "<image_soft_token>",
|
| 19 |
+
"pad_token": {
|
| 20 |
+
"content": "<pad>",
|
| 21 |
+
"lstrip": false,
|
| 22 |
+
"normalized": false,
|
| 23 |
+
"rstrip": false,
|
| 24 |
+
"single_word": false
|
| 25 |
+
},
|
| 26 |
+
"unk_token": {
|
| 27 |
+
"content": "<unk>",
|
| 28 |
+
"lstrip": false,
|
| 29 |
+
"normalized": false,
|
| 30 |
+
"rstrip": false,
|
| 31 |
+
"single_word": false
|
| 32 |
+
}
|
| 33 |
+
}
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:216e2a79606fe879c9f17c529c71cd241338407fd5646b595ffd3c4b9ea1d503
|
| 3 |
+
size 33385262
|
tokenizer_config.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
vulkan-derivation.json
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"artifacts": [
|
| 3 |
+
{
|
| 4 |
+
"bytes": 4392411,
|
| 5 |
+
"path": "model.onnx",
|
| 6 |
+
"sha256": "04eb85c8ab0991122030856f246f4f9cb922a7c61c3045b8978a32df5d1cb4d3"
|
| 7 |
+
},
|
| 8 |
+
{
|
| 9 |
+
"bytes": 1230372864,
|
| 10 |
+
"path": "model.onnx.data",
|
| 11 |
+
"sha256": "56af931663f87a32abce8dd213bc87e595c1aafb12668b60408bf005cace2247"
|
| 12 |
+
}
|
| 13 |
+
],
|
| 14 |
+
"changes": [
|
| 15 |
+
{
|
| 16 |
+
"dtype": 7,
|
| 17 |
+
"name": "node_abs_1",
|
| 18 |
+
"operation": "Abs"
|
| 19 |
+
}
|
| 20 |
+
],
|
| 21 |
+
"max_sequence": 1024,
|
| 22 |
+
"role": "gemma",
|
| 23 |
+
"schema": "daecore.vulkan-mask-derivation.v1",
|
| 24 |
+
"source_graph_sha256": "07dec7832d2c15c498c059c17359da9335e746ad5270d1cd4665a06789e388f1",
|
| 25 |
+
"source_weights_sha256": "56af931663f87a32abce8dd213bc87e595c1aafb12668b60408bf005cace2247",
|
| 26 |
+
"weights_unchanged": true
|
| 27 |
+
}
|