tnh0527's picture
Publish embeddinggemma-300m-memory-ft-v2 (qualified ONNX export and complete attribution)
bf5a9d3 verified
|
Raw History Blame Contribute Delete
2.74 kB

Modifications from the upstream model

embeddinggemma-300m-memory-ft-v2 is not an unmodified copy of google/embeddinggemma-300m. It is a Daecore-modified derivative and is not endorsed by Google.

What changed:

  • Fine-tune. The base weights were trained further for semantic retrieval over Daecore's governed memory corpus, using query and document role prefixes, graded positive supervision, hard negatives, Matryoshka dimensions, and attention LoRA (rank 16, alpha 32). The adapter was merged into the model weights. Relevance grades were frontier-model judgments under a frozen protocol, not human annotations.
  • Selected successor. A fresh upstream fit combined cleaned private supervision with annotated HotpotQA, MultiDoc2Dial and FinQA evidence, upstream retention (weight 2), and a within-question grade-3 preference (weight 0.25). The selected merged weights are from update 1,130. The model card records their actual data exposure and measured tradeoffs; the prior model's counts and qualification do not describe these weights.
  • Export. The merged model was exported to ONNX as one graph taking input_ids and attention_mask and emitting 768-dimensional embeddings in FP32, with the parameters stored as external data.
  • Serving contract. serving.json records the runtime and sequence geometry of the original export qualification: 128 query tokens, 1,024 passage tokens and a CUDA batch of 32. Production batching adapts to available memory within the separately qualified CPU, CUDA and Vulkan limits. Publication and source activation remain separate from a successful export.

Files modified or generated by Daecore:

  • model.onnx: the generated ONNX graph of the fine-tuned model;
  • model.onnx.data: the fine-tuned parameters used by that graph;
  • config.json: the export configuration of the modified model;
  • serving.json: the Daecore runtime and sequence-geometry contract.
  • vulkan-derivation.json: exact source and derived graph identities and the bounded mask rewrite, with learned weights unchanged by that rewrite.

Unchanged: tokenizer.json, tokenizer_config.json, special_tokens_map.json, and added_tokens.json are the tokenizer files of the selected model export and bind the exact text-processing contract.

This distribution is subject to the Gemma Terms of Use (LICENSE) and the Gemma Prohibited Use Policy (GEMMA_PROHIBITED_USE_POLICY.txt).

  • Portable shape operations. Inferred-dimension reshape operators use allowzero=0. A bounded integer mask absolute-value operation runs in exact FP32 and casts back to its original type, allowing native Vulkan execution. CPU, CUDA and Vulkan use the same derived graph.