|
Download MODIFICATIONS.md from Daecore/embeddinggemma-300m-memory-ft-v2: direct link, hf CLI and curl.
- Browser
- Download file 2.74 kB
-
https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2/resolve/766fddfa2542ca5735aba9fa9959a19d5c38598e/MODIFICATIONS.md
- Command line
-
hf download hf://Daecore/embeddinggemma-300m-memory-ft-v2@766fddfa2542ca5735aba9fa9959a19d5c38598e/MODIFICATIONS.md
-
curl -L -o MODIFICATIONS.md https://huggingface.co/Daecore/embeddinggemma-300m-memory-ft-v2/resolve/766fddfa2542ca5735aba9fa9959a19d5c38598e/MODIFICATIONS.md
2.74 kB
| # Modifications from the upstream model | |
| `embeddinggemma-300m-memory-ft-v2` is not an unmodified copy of | |
| `google/embeddinggemma-300m`. It is a Daecore-modified derivative and is not | |
| endorsed by Google. | |
| What changed: | |
| - **Fine-tune.** The base weights were trained further for semantic retrieval | |
| over Daecore's governed memory corpus, using query and document role | |
| prefixes, graded positive supervision, hard negatives, Matryoshka | |
| dimensions, and attention LoRA (rank 16, alpha 32). The adapter was merged | |
| into the model weights. Relevance grades were frontier-model judgments under | |
| a frozen protocol, not human annotations. | |
| - **Selected successor.** A fresh upstream fit combined cleaned private | |
| supervision with annotated HotpotQA, MultiDoc2Dial and FinQA evidence, | |
| upstream retention (weight 2), and a within-question grade-3 preference | |
| (weight 0.25). The selected merged weights are from update 1,130. The model | |
| card records their actual data exposure and measured tradeoffs; the prior | |
| model's counts and qualification do not describe these weights. | |
| - **Export.** The merged model was exported to ONNX as one graph taking | |
| `input_ids` and `attention_mask` and emitting 768-dimensional embeddings in | |
| FP32, with the parameters stored as external data. | |
| - **Serving contract.** `serving.json` records the runtime and sequence | |
| geometry of the original export qualification: 128 query tokens, 1,024 passage | |
| tokens and a CUDA batch of 32. Production batching adapts to available memory | |
| within the separately qualified CPU, CUDA and Vulkan limits. Publication | |
| and source activation remain separate from a successful export. | |
| Files modified or generated by Daecore: | |
| - `model.onnx`: the generated ONNX graph of the fine-tuned model; | |
| - `model.onnx.data`: the fine-tuned parameters used by that graph; | |
| - `config.json`: the export configuration of the modified model; | |
| - `serving.json`: the Daecore runtime and sequence-geometry contract. | |
| - `vulkan-derivation.json`: exact source and derived graph identities and the | |
| bounded mask rewrite, with learned weights unchanged by that rewrite. | |
| Unchanged: `tokenizer.json`, `tokenizer_config.json`, | |
| `special_tokens_map.json`, and `added_tokens.json` are the tokenizer files of | |
| the selected model export and bind the exact text-processing contract. | |
| This distribution is subject to the Gemma Terms of Use (`LICENSE`) and the | |
| Gemma Prohibited Use Policy (`GEMMA_PROHIBITED_USE_POLICY.txt`). | |
| - **Portable shape operations.** Inferred-dimension reshape operators use | |
| `allowzero=0`. A bounded integer mask absolute-value operation runs in exact | |
| FP32 and casts back to its original type, allowing native Vulkan execution. | |
| CPU, CUDA and Vulkan use the same derived graph. | |