File size: 2,739 Bytes
bf5a9d3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 | # Modifications from the upstream model
`embeddinggemma-300m-memory-ft-v2` is not an unmodified copy of
`google/embeddinggemma-300m`. It is a Daecore-modified derivative and is not
endorsed by Google.
What changed:
- **Fine-tune.** The base weights were trained further for semantic retrieval
over Daecore's governed memory corpus, using query and document role
prefixes, graded positive supervision, hard negatives, Matryoshka
dimensions, and attention LoRA (rank 16, alpha 32). The adapter was merged
into the model weights. Relevance grades were frontier-model judgments under
a frozen protocol, not human annotations.
- **Selected successor.** A fresh upstream fit combined cleaned private
supervision with annotated HotpotQA, MultiDoc2Dial and FinQA evidence,
upstream retention (weight 2), and a within-question grade-3 preference
(weight 0.25). The selected merged weights are from update 1,130. The model
card records their actual data exposure and measured tradeoffs; the prior
model's counts and qualification do not describe these weights.
- **Export.** The merged model was exported to ONNX as one graph taking
`input_ids` and `attention_mask` and emitting 768-dimensional embeddings in
FP32, with the parameters stored as external data.
- **Serving contract.** `serving.json` records the runtime and sequence
geometry of the original export qualification: 128 query tokens, 1,024 passage
tokens and a CUDA batch of 32. Production batching adapts to available memory
within the separately qualified CPU, CUDA and Vulkan limits. Publication
and source activation remain separate from a successful export.
Files modified or generated by Daecore:
- `model.onnx`: the generated ONNX graph of the fine-tuned model;
- `model.onnx.data`: the fine-tuned parameters used by that graph;
- `config.json`: the export configuration of the modified model;
- `serving.json`: the Daecore runtime and sequence-geometry contract.
- `vulkan-derivation.json`: exact source and derived graph identities and the
bounded mask rewrite, with learned weights unchanged by that rewrite.
Unchanged: `tokenizer.json`, `tokenizer_config.json`,
`special_tokens_map.json`, and `added_tokens.json` are the tokenizer files of
the selected model export and bind the exact text-processing contract.
This distribution is subject to the Gemma Terms of Use (`LICENSE`) and the
Gemma Prohibited Use Policy (`GEMMA_PROHIBITED_USE_POLICY.txt`).
- **Portable shape operations.** Inferred-dimension reshape operators use
`allowzero=0`. A bounded integer mask absolute-value operation runs in exact
FP32 and casts back to its original type, allowing native Vulkan execution.
CPU, CUDA and Vulkan use the same derived graph.
|