Text Generation
PEFT
TensorBoard
Safetensors
Transformers
English
lora
kappaTune
conversational
oswaldoludwig commited on
Commit
addac3c
·
verified ·
1 Parent(s): a661295

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +3 -3
README.md CHANGED
@@ -31,9 +31,9 @@ This model is a fine-tuned version of [TinyLlama/TinyLlama-1.1B-Chat-v1.0](https
31
  ## Model description
32
 
33
  This model is a LoRA (Low-Rank Adaptation) adapter applied to TinyLlama-1.1B. Unlike standard LoRA, which targets manually specified module types (e.g., all `q_proj` or `v_proj` layers), this adapter was trained using the **KappaTune PEFT Integration**.
34
- Before training, the KappaTune algorithm performed a Singular Value Decomposition (SVD) on all candidate weight matrices to calculate their **Condition Number (\( \kappa \))**.
35
- * **High-\( \kappa \)** tensors (highly specialized, anisotropic weights containing pre-trained knowledge) were frozen.
36
- * **Low-\( \kappa \)** tensors (numerically stable, general-purpose weights acting as a "raw marble block") were selected for LoRA adaptation.
37
 
38
  This targeted approach ensures the model learns the new domain efficiently while preserving its foundational conversational capabilities.
39
 
 
31
  ## Model description
32
 
33
  This model is a LoRA (Low-Rank Adaptation) adapter applied to TinyLlama-1.1B. Unlike standard LoRA, which targets manually specified module types (e.g., all `q_proj` or `v_proj` layers), this adapter was trained using the **KappaTune PEFT Integration**.
34
+ Before training, the KappaTune algorithm performed a Singular Value Decomposition (SVD) on all candidate weight matrices to calculate their **Condition Number (kappa)**.
35
+ * **High-kappa** tensors (highly specialized, anisotropic weights containing pre-trained knowledge) were frozen.
36
+ * **Low-kappa** tensors (numerically stable, general-purpose weights of higher output entropy acting as a "raw marble block") were selected for LoRA adaptation.
37
 
38
  This targeted approach ensures the model learns the new domain efficiently while preserving its foundational conversational capabilities.
39