Instructions to use hoin1218/bge-m3-merchant-region-pair-adapter with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use hoin1218/bge-m3-merchant-region-pair-adapter with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("hoin1218/bge-m3-merchant-region-pair-adapter") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
BGE-M3 Merchant Region Pair Similarity Adapter
hoin1218/bge-m3-business-compact-card-context-v2 ๊ธฐ๋ฐ์ ๊ฐ๋งน์ pair ์ ์ฌ๋ adapter์
๋๋ค.
Interactive demo: https://huggingface.co/spaces/hoin1218/bge-m3-merchant-region-pair-demo
์
๋ ฅ์ ์๋์ฒ๋ผ ๊ฐ๋งน์ ๋ช
+ ์
์ข
๋ช
+ ์ง์ญ ๋ ๊ฐ๊ฐ ๋ค์ด์ค๋ ์ํฉ์ ๊ฐ์ ํ์ต๋๋ค.
๊ฐ๋งน์ ๋ช
: ์คํ๋ฒ
์ค ๊ฐ๋จ์ญ์ | ์
์ข
๋ช
: ์ปคํผ์ ๋ฌธ์ | ์ง์ญ: ์์ธ ๊ฐ๋จ๊ตฌ
๊ฐ๋งน์ ๋ช
: ์ด๋์ผ์ปคํผ ์ญ์ผ์ | ์
์ข
๋ช
: ์ปคํผ์ ๋ฌธ์ | ์ง์ญ: ์์ธ ๊ฐ๋จ๊ตฌ
Current Version
adapter/์๋ ์ต์ adapter๋ฅผ ๋ฃ์์ต๋๋ค.
์ต์ adapter๋ ๋ค์ ์์๋ก ํ์ต๋์์ต๋๋ค.
- synthetic merchant pair dataset์ผ๋ก projection adapter ํ์ต
- ์ค์ ์๊ฐ์ ๋ณด row ๊ธฐ๋ฐ pair dataset์ผ๋ก ์ถ๊ฐ ํ์ต
์ด์ synthetic-only adapter๋ synthetic_adapter/์ ๋ณด๊ดํ์ต๋๋ค. ์ต์ real-data continued adapter๋ real_adapter/์๋ ๋์ผํ๊ฒ ๋ค์ด ์์ต๋๋ค.
Included Files
adapter/: latest real-data continued adaptersynthetic_adapter/: previous synthetic-only adapterreal_adapter/: latest real-data continued adapter copydata/: synthetic pair datasetreal_data/: real merchant row based pair datasetresults/: synthetic pair evaluationreal_results/: real pair evaluationreal_results/judgement_check_ko.md: manual judgement check summaryscripts/: download/build/train/evaluation scriptsDATASETS_USED.md: used and candidate dataset details
Real Merchant Pair Dataset
์ค์ ํ์ต์ ์ถ๊ฐ๋ก ์ฌ์ฉํ ๋ฐ์ดํฐ๋ ์์๊ณต์ธ์์ฅ์งํฅ๊ณต๋จ_์๊ฐ(์๊ถ)์ ๋ณด_202506 ๊ณต๊ฐ mirror์์ ์ผ๋ถ ๊ถ์ญ์ ๋ฐ์ pair๋ก ๋ณํํ ๋ฐ์ดํฐ์
๋๋ค.
Raw source mirror:
- Hugging Face dataset:
ginipick/market - Original source:
์์๊ณต์ธ์์ฅ์งํฅ๊ณต๋จ_์๊ฐ(์๊ถ)์ ๋ณด - Regions used:
- ์ธ์ข
- ์ธ์ฐ
- ์ ์ฃผ
- ๊ด์ฃผ
- ๋์
Raw CSV ์ ์ฒด๋ฅผ ์ด repo์ ๋ค์ ํฌํจํ์ง๋ ์์์ต๋๋ค. ๋์ ํ์ต์ ์ฌ์ฉํ pair ๋ฐ์ดํฐ์ source manifest๋ฅผ ํฌํจํ์ต๋๋ค.
Real pair dataset:
- sampled source records: 53,748
- total pair rows: 24,500
- train: 19,600
- eval: 2,450
- test: 2,450
Pair label policy:
| Pair ์ ํ | Score |
|---|---|
| ๊ฐ์ ๊ฐ๋งน์ / ๊ฐ์ ์ ์ข / ๊ฐ์ ์ง์ญ | 1.00 |
| ๊ฐ์ ์ํธ๋ช / ๊ฐ์ ์ ์ข / ๋ค๋ฅธ ์ง์ญ | 0.72 |
| ๊ฐ์ ์๋ถ๋ฅ ์ ์ข / ๊ฐ์ ์ง์ญ | 0.82 |
| ๊ฐ์ ์๋ถ๋ฅ ์ ์ข / ๋ค๋ฅธ ์ง์ญ | 0.58 |
| ๊ฐ์ ์ค๋ถ๋ฅ ์ ์ข / ๋ค๋ฅธ ์๋ถ๋ฅ / ๊ฐ์ ์ง์ญ | 0.35 |
| ๋ค๋ฅธ ๋๋ถ๋ฅ ์ ์ข / ๊ฐ์ ์ง์ญ | 0.10 |
| ๋ค๋ฅธ ๋๋ถ๋ฅ ์ ์ข / ๋ค๋ฅธ ์ง์ญ | 0.02 |
Evaluation On Real Pair Test Set
| Model | Pearson | Spearman | MAE | RMSE |
|---|---|---|---|---|
| previous_synthetic_pair_adapter | 0.9156 | 0.9033 | 0.1720 | 0.2228 |
| real_pair_continued_adapter | 0.9257 | 0.9081 | 0.1569 | 0.2082 |
Mean score by reason:
| Reason | Previous synthetic adapter | Real-data continued adapter |
|---|---|---|
| different_large_industry_different_region | 0.1141 | 0.0549 |
| different_large_industry_same_region | 0.1152 | 0.0580 |
| same_mid_industry_different_small_same_region | 0.5015 | 0.4290 |
| same_name_same_industry_different_region | 0.9906 | 0.9884 |
| same_small_industry_different_region | 0.9719 | 0.9674 |
| same_small_industry_same_region | 0.9800 | 0.9771 |
Manual Check
Detailed Korean judgement report: real_results/judgement_check_ko.md
Manual examples using older synthetic labels:
| Pair | Synthetic adapter | Real-data continued adapter |
|---|---|---|
| ์ปคํผ์ ๋ฌธ์ vs ์ปคํผ์ ๋ฌธ์ , ๊ฐ์ ์ง์ญ | 0.9226 | 0.9428 |
| ์ปคํผ์ ๋ฌธ์ vs ์ปคํผ์ ๋ฌธ์ , ๋ค๋ฅธ ์ง์ญ | 0.8614 | 0.8948 |
| ์ปคํผ์ ๋ฌธ์ vs ์ ๊ณผ์ , ๊ฐ์ ์ง์ญ | 0.0499 | 0.2866 |
| ์ปคํผ์ ๋ฌธ์ vs ์ฃผ์ ์, ๊ฐ์ ์ง์ญ | 0.0893 | 0.1526 |
| ์ปคํผ์ ๋ฌธ์ vs ์ฃผ์ ์, ๋ค๋ฅธ ์ง์ญ | 0.0259 | 0.0798 |
The real-data continued adapter improves the real pair test set, but it is tuned to the public store dataset's industry taxonomy. If production input uses labels like ์ปคํผ์ ๋ฌธ์ , ์ฃผ์ ์, or card-company industry names, those labels should be included in the next training pass.
Usage
from scripts.load_merchant_region_pair_adapter import load_pair_similarity_model, similarity
model = load_pair_similarity_model("adapter")
text1 = "๊ฐ๋งน์ ๋ช
: ์คํ๋ฒ
์ค ๊ฐ๋จ์ญ์ | ์
์ข
๋ช
: ์ปคํผ์ ๋ฌธ์ | ์ง์ญ: ์์ธ ๊ฐ๋จ๊ตฌ"
text2 = "๊ฐ๋งน์ ๋ช
: ์ด๋์ผ์ปคํผ ์ญ์ผ์ | ์
์ข
๋ช
: ์ปคํผ์ ๋ฌธ์ | ์ง์ญ: ์์ธ ๊ฐ๋จ๊ตฌ"
print(similarity(model, text1, text2))
Limitation
This is still a proof-of-concept adapter. The new training data uses real merchant rows, but pair scores are automatically generated from industry/region rules. For production, add labeled or policy-approved pairs from the actual card merchant taxonomy.