Yuki131 commited on
Commit
880719b
·
verified ·
1 Parent(s): a96ec65

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -4
README.md CHANGED
@@ -1,15 +1,12 @@
1
  ---
2
  license: apache-2.0
3
  base_model: Qwen/Qwen3-Embedding-4B
4
- base_model_relation: merge
5
  datasets:
6
  - HIT-TMG/JevEmbed-Data
7
  pipeline_tag: feature-extraction
8
  library_name: sentence-transformers
9
  tags:
10
- - sentence-transformers
11
  - jevembed
12
- - lora
13
  ---
14
 
15
  # JevEmbed-Qwen3-Embedding-4B
@@ -44,4 +41,4 @@ Qwen3-Embedding-4B and the final JevEmbed LoRA checkpoint were evaluated on all
44
 
45
  Training ran for one epoch on the 1,601,157-question JevEmbed-Data training split. The final checkpoint is step **3,127**. Training used 16 GPUs across four nodes, per-GPU batch 4, gradient accumulation 8 (effective batch 512), BF16, LoRA rank 64, alpha 32, dropout 0.05, and Q/K/V projection targets. The learning rate was 2 × 10⁻⁴ with 10% warmup. Inputs were truncated at 1,024 tokens. The training objective is described in the [JevEmbed implementation](https://github.com/HITsz-TMG/JevEmbed/blob/main/src/jevembed/training/objective.py).
46
 
47
- The source Qwen model is Apache 2.0 licensed. Training-data source licenses vary; see the dataset's [source report](https://huggingface.co/datasets/HIT-TMG/JevEmbed-Data/blob/main/docs/PROCESSING_REPORT.md).
 
1
  ---
2
  license: apache-2.0
3
  base_model: Qwen/Qwen3-Embedding-4B
 
4
  datasets:
5
  - HIT-TMG/JevEmbed-Data
6
  pipeline_tag: feature-extraction
7
  library_name: sentence-transformers
8
  tags:
 
9
  - jevembed
 
10
  ---
11
 
12
  # JevEmbed-Qwen3-Embedding-4B
 
41
 
42
  Training ran for one epoch on the 1,601,157-question JevEmbed-Data training split. The final checkpoint is step **3,127**. Training used 16 GPUs across four nodes, per-GPU batch 4, gradient accumulation 8 (effective batch 512), BF16, LoRA rank 64, alpha 32, dropout 0.05, and Q/K/V projection targets. The learning rate was 2 × 10⁻⁴ with 10% warmup. Inputs were truncated at 1,024 tokens. The training objective is described in the [JevEmbed implementation](https://github.com/HITsz-TMG/JevEmbed/blob/main/src/jevembed/training/objective.py).
43
 
44
+ The source Qwen model is Apache 2.0 licensed. Training-data source licenses vary; see the dataset's [source report](https://huggingface.co/datasets/HIT-TMG/JevEmbed-Data/blob/main/docs/PROCESSING_REPORT.md).