How to use from the
Use from the
sentence-transformers library
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("hoin1218/bge-m3-merchant-region-pair-adapter")

sentences = [
    "That is a happy person",
    "That is a happy dog",
    "That is a very happy person",
    "Today is a sunny day"
]
embeddings = model.encode(sentences)

similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)
# [4, 4]

BGE-M3 Merchant Region Pair Similarity Adapter

hoin1218/bge-m3-business-compact-card-context-v2 ๊ธฐ๋ฐ˜์˜ ๊ฐ€๋งน์  pair ์œ ์‚ฌ๋„ adapter์ž…๋‹ˆ๋‹ค.

Interactive demo: https://huggingface.co/spaces/hoin1218/bge-m3-merchant-region-pair-demo

์ž…๋ ฅ์€ ์•„๋ž˜์ฒ˜๋Ÿผ ๊ฐ€๋งน์ ๋ช… + ์—…์ข…๋ช… + ์ง€์—ญ ๋‘ ๊ฐœ๊ฐ€ ๋“ค์–ด์˜ค๋Š” ์ƒํ™ฉ์„ ๊ฐ€์ •ํ–ˆ์Šต๋‹ˆ๋‹ค.

๊ฐ€๋งน์ ๋ช…: ์Šคํƒ€๋ฒ…์Šค ๊ฐ•๋‚จ์—ญ์  | ์—…์ข…๋ช…: ์ปคํ”ผ์ „๋ฌธ์  | ์ง€์—ญ: ์„œ์šธ ๊ฐ•๋‚จ๊ตฌ
๊ฐ€๋งน์ ๋ช…: ์ด๋””์•ผ์ปคํ”ผ ์—ญ์‚ผ์  | ์—…์ข…๋ช…: ์ปคํ”ผ์ „๋ฌธ์  | ์ง€์—ญ: ์„œ์šธ ๊ฐ•๋‚จ๊ตฌ

Current Version

adapter/์—๋Š” ์ตœ์‹  adapter๋ฅผ ๋„ฃ์—ˆ์Šต๋‹ˆ๋‹ค.

์ตœ์‹  adapter๋Š” ๋‹ค์Œ ์ˆœ์„œ๋กœ ํ•™์Šต๋˜์—ˆ์Šต๋‹ˆ๋‹ค.

  1. synthetic merchant pair dataset์œผ๋กœ projection adapter ํ•™์Šต
  2. ์‹ค์ œ ์ƒ๊ฐ€์ •๋ณด row ๊ธฐ๋ฐ˜ pair dataset์œผ๋กœ ์ถ”๊ฐ€ ํ•™์Šต

์ด์ „ synthetic-only adapter๋Š” synthetic_adapter/์— ๋ณด๊ด€ํ–ˆ์Šต๋‹ˆ๋‹ค. ์ตœ์‹  real-data continued adapter๋Š” real_adapter/์—๋„ ๋™์ผํ•˜๊ฒŒ ๋“ค์–ด ์žˆ์Šต๋‹ˆ๋‹ค.

Included Files

  • adapter/: latest real-data continued adapter
  • synthetic_adapter/: previous synthetic-only adapter
  • real_adapter/: latest real-data continued adapter copy
  • data/: synthetic pair dataset
  • real_data/: real merchant row based pair dataset
  • results/: synthetic pair evaluation
  • real_results/: real pair evaluation
  • real_results/judgement_check_ko.md: manual judgement check summary
  • scripts/: download/build/train/evaluation scripts
  • DATASETS_USED.md: used and candidate dataset details

Real Merchant Pair Dataset

์‹ค์ œ ํ•™์Šต์— ์ถ”๊ฐ€๋กœ ์‚ฌ์šฉํ•œ ๋ฐ์ดํ„ฐ๋Š” ์†Œ์ƒ๊ณต์ธ์‹œ์žฅ์ง„ํฅ๊ณต๋‹จ_์ƒ๊ฐ€(์ƒ๊ถŒ)์ •๋ณด_202506 ๊ณต๊ฐœ mirror์—์„œ ์ผ๋ถ€ ๊ถŒ์—ญ์„ ๋ฐ›์•„ pair๋กœ ๋ณ€ํ™˜ํ•œ ๋ฐ์ดํ„ฐ์ž…๋‹ˆ๋‹ค.

Raw source mirror:

  • Hugging Face dataset: ginipick/market
  • Original source: ์†Œ์ƒ๊ณต์ธ์‹œ์žฅ์ง„ํฅ๊ณต๋‹จ_์ƒ๊ฐ€(์ƒ๊ถŒ)์ •๋ณด
  • Regions used:
    • ์„ธ์ข…
    • ์šธ์‚ฐ
    • ์ œ์ฃผ
    • ๊ด‘์ฃผ
    • ๋Œ€์ „

Raw CSV ์ „์ฒด๋ฅผ ์ด repo์— ๋‹ค์‹œ ํฌํ•จํ•˜์ง€๋Š” ์•Š์•˜์Šต๋‹ˆ๋‹ค. ๋Œ€์‹  ํ•™์Šต์— ์‚ฌ์šฉํ•œ pair ๋ฐ์ดํ„ฐ์™€ source manifest๋ฅผ ํฌํ•จํ–ˆ์Šต๋‹ˆ๋‹ค.

Real pair dataset:

  • sampled source records: 53,748
  • total pair rows: 24,500
  • train: 19,600
  • eval: 2,450
  • test: 2,450

Pair label policy:

Pair ์œ ํ˜• Score
๊ฐ™์€ ๊ฐ€๋งน์  / ๊ฐ™์€ ์—…์ข… / ๊ฐ™์€ ์ง€์—ญ 1.00
๊ฐ™์€ ์ƒํ˜ธ๋ช… / ๊ฐ™์€ ์—…์ข… / ๋‹ค๋ฅธ ์ง€์—ญ 0.72
๊ฐ™์€ ์†Œ๋ถ„๋ฅ˜ ์—…์ข… / ๊ฐ™์€ ์ง€์—ญ 0.82
๊ฐ™์€ ์†Œ๋ถ„๋ฅ˜ ์—…์ข… / ๋‹ค๋ฅธ ์ง€์—ญ 0.58
๊ฐ™์€ ์ค‘๋ถ„๋ฅ˜ ์—…์ข… / ๋‹ค๋ฅธ ์†Œ๋ถ„๋ฅ˜ / ๊ฐ™์€ ์ง€์—ญ 0.35
๋‹ค๋ฅธ ๋Œ€๋ถ„๋ฅ˜ ์—…์ข… / ๊ฐ™์€ ์ง€์—ญ 0.10
๋‹ค๋ฅธ ๋Œ€๋ถ„๋ฅ˜ ์—…์ข… / ๋‹ค๋ฅธ ์ง€์—ญ 0.02

Evaluation On Real Pair Test Set

Model Pearson Spearman MAE RMSE
previous_synthetic_pair_adapter 0.9156 0.9033 0.1720 0.2228
real_pair_continued_adapter 0.9257 0.9081 0.1569 0.2082

Mean score by reason:

Reason Previous synthetic adapter Real-data continued adapter
different_large_industry_different_region 0.1141 0.0549
different_large_industry_same_region 0.1152 0.0580
same_mid_industry_different_small_same_region 0.5015 0.4290
same_name_same_industry_different_region 0.9906 0.9884
same_small_industry_different_region 0.9719 0.9674
same_small_industry_same_region 0.9800 0.9771

Manual Check

Detailed Korean judgement report: real_results/judgement_check_ko.md

Manual examples using older synthetic labels:

Pair Synthetic adapter Real-data continued adapter
์ปคํ”ผ์ „๋ฌธ์  vs ์ปคํ”ผ์ „๋ฌธ์ , ๊ฐ™์€ ์ง€์—ญ 0.9226 0.9428
์ปคํ”ผ์ „๋ฌธ์  vs ์ปคํ”ผ์ „๋ฌธ์ , ๋‹ค๋ฅธ ์ง€์—ญ 0.8614 0.8948
์ปคํ”ผ์ „๋ฌธ์  vs ์ œ๊ณผ์ , ๊ฐ™์€ ์ง€์—ญ 0.0499 0.2866
์ปคํ”ผ์ „๋ฌธ์  vs ์ฃผ์œ ์†Œ, ๊ฐ™์€ ์ง€์—ญ 0.0893 0.1526
์ปคํ”ผ์ „๋ฌธ์  vs ์ฃผ์œ ์†Œ, ๋‹ค๋ฅธ ์ง€์—ญ 0.0259 0.0798

The real-data continued adapter improves the real pair test set, but it is tuned to the public store dataset's industry taxonomy. If production input uses labels like ์ปคํ”ผ์ „๋ฌธ์ , ์ฃผ์œ ์†Œ, or card-company industry names, those labels should be included in the next training pass.

Usage

from scripts.load_merchant_region_pair_adapter import load_pair_similarity_model, similarity

model = load_pair_similarity_model("adapter")

text1 = "๊ฐ€๋งน์ ๋ช…: ์Šคํƒ€๋ฒ…์Šค ๊ฐ•๋‚จ์—ญ์  | ์—…์ข…๋ช…: ์ปคํ”ผ์ „๋ฌธ์  | ์ง€์—ญ: ์„œ์šธ ๊ฐ•๋‚จ๊ตฌ"
text2 = "๊ฐ€๋งน์ ๋ช…: ์ด๋””์•ผ์ปคํ”ผ ์—ญ์‚ผ์  | ์—…์ข…๋ช…: ์ปคํ”ผ์ „๋ฌธ์  | ์ง€์—ญ: ์„œ์šธ ๊ฐ•๋‚จ๊ตฌ"

print(similarity(model, text1, text2))

Limitation

This is still a proof-of-concept adapter. The new training data uses real merchant rows, but pair scores are automatically generated from industry/region rules. For production, add labeled or policy-approved pairs from the actual card merchant taxonomy.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for hoin1218/bge-m3-merchant-region-pair-adapter

Finetuned
(1)
this model

Space using hoin1218/bge-m3-merchant-region-pair-adapter 1