--- license: apache-2.0 base_model: Qwen/Qwen3-Embedding-4B datasets: - HIT-TMG/JevEmbed-Data pipeline_tag: feature-extraction library_name: sentence-transformers tags: - jevembed --- # JevEmbed-Qwen3-Embedding-4B JevEmbed-Qwen3-Embedding-4B is a fine-tuned version of [Qwen3-Embedding-4B](https://huggingface.co/Qwen/Qwen3-Embedding-4B) for [JevEmbed](https://github.com/HITsz-TMG/JevEmbed) Choice, Score, and Noul decisions. The LoRA adapter is merged into a standalone Sentence Transformers model; the adapter is also available in [`lora/`](lora/). Embeddings have 2,560 dimensions and use last-token pooling and normalization. Tested with `transformers==4.51.0` and `sentence-transformers==5.3.0`. ## Use with JevEmbed Install JevEmbed with its local-model dependencies, download [`jevembed.yaml`](jevembed.yaml), and pass your own request JSON to the CLI: ```bash python -m pip install 'jevembed[local] @ git+https://github.com/HITsz-TMG/JevEmbed.git' python -m jevembed --config jevembed.yaml --input /path/to/your/request.json ``` The weights download automatically from `HIT-TMG/JevEmbed-Qwen3-Embedding-4B`. If the request has a `model` field, use `jevembed-qwen3-embedding-4b` or the full Hugging Face ID. To use weights already on disk, set `model_name_or_path` to that directory and `local_files_only: true` in the YAML. The configuration uses 1,024-token truncation, Choice/Score temperature 0.1, and Noul slope 10. Raw embeddings can be loaded with `SentenceTransformer("HIT-TMG/JevEmbed-Qwen3-Embedding-4B")`; JevEmbed applies the task prompts and scoring. ## Test-set performance Qwen3-Embedding-4B and the final JevEmbed LoRA checkpoint were evaluated on all 66,482 questions in the [JevEmbed-Data](https://huggingface.co/datasets/HIT-TMG/JevEmbed-Data) test split. These results are from the final LoRA checkpoint before merging. Both runs used BF16, identical JevEmbed prompts and scoring, and 1,024-token truncation. Accuracy uses hard labels; MAE also includes soft targets where present. | Metric | Base | Final LoRA checkpoint | Change | | --- | ---: | ---: | ---: | | Overall hard-label accuracy (64,110) | 36.29% | 85.86% | +49.57 pp | | Choice accuracy (17,487) | 38.11% | 90.31% | +52.20 pp | | Score level accuracy (24,260) | 30.55% | 73.41% | +42.86 pp | | Noul binary accuracy (22,363) | 41.09% | 95.88% | +54.78 pp | | Score MAE (24,287; lower is better) | 0.9211 | 0.3628 | -0.5583 | | Noul MAE (24,004; lower is better) | 0.5895 | 0.0689 | -0.5206 | ## Training Training ran for one epoch on the 1,601,157-question JevEmbed-Data training split. The final checkpoint is step **3,127**. Training used 16 GPUs across four nodes, per-GPU batch 4, gradient accumulation 8 (effective batch 512), BF16, LoRA rank 64, alpha 32, dropout 0.05, and Q/K/V projection targets. The learning rate was 2 × 10⁻⁴ with 10% warmup. Inputs were truncated at 1,024 tokens. The training objective is described in the [JevEmbed implementation](https://github.com/HITsz-TMG/JevEmbed/blob/main/src/jevembed/training/objective.py). The source Qwen model is Apache 2.0 licensed. Training-data source licenses vary; see the dataset's [source report](https://huggingface.co/datasets/HIT-TMG/JevEmbed-Data/blob/main/docs/PROCESSING_REPORT.md).