Spaces:
Running
Running
File size: 3,631 Bytes
68f5f5e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 | # Benchmark λ°μ΄ν°μ
λ°°ν¬ κ΅¬μ±
Hugging Face Dataset μ μ₯μλ₯Ό μλ λ¨μλ‘ κ΄λ¦¬ν©λλ€. κ°μΈ κ³μ μμ λ¨Όμ λ°°ν¬νκ³
κ²μ¦ν λ€, νμνλ©΄ κ° μ μ₯μλ₯Ό `KETI-NLP` μ‘°μ§μΌλ‘ μ΄μ ν©λλ€.
μλ μ μ₯μ IDμ κ²½λ‘λ λ°°ν¬ν λ μ¬μ©ν μ μμ΄λ©°, μμ§ μμ±λ μ μ₯μλ₯Ό μλ―Ένμ§ μμ΅λλ€.
| νμ μ΄λ¦ | κ°μΈ μ μ₯μ ID μ μ | ν¬ν¨ λ°μ΄ν°μ
|
| --- | --- | --- |
| μ§μ€μ± Benchmark | `DoolyKim22/Truthfulness-Benchmark` | μΌλ° μμ, μμΉ μ 보, ν©νΈ 체ν¬, λ©ν° λͺ¨λ¬ |
| μλ²λ¦° Benchmark | `DoolyKim22/Sovereign-Benchmark` | μλ²λ¦° |
| K-Prism | `DoolyKim22/K-Prism` | K-Prism Text / Image νΈλ |
μ§μ€μ±μ λ€ νλͺ©μ νλμ μ μ₯μ μμμ λ³λμ configurationμΌλ‘ μ 곡ν©λλ€.
κ° νλͺ©μ μλ³Έ νλμ μ λ΅ μ²΄κ³λ μ μ§νλ©°, νλμ JSONμΌλ‘ ν©μΉμ§ μμ΅λλ€.
μ΄ κ΅¬μ±μ λ°μ΄ν° λ°°ν¬ λ¨μμ΄λ©°, 리λ보λμ νκ° νλͺ©μ΄λ νκ· κ³μ°μ λ³κ²½νμ§ μμ΅λλ€.
## μ§μ€μ± Benchmark
λ€μμ JSON νμμ νκ° λ°μ΄ν°λ₯Ό μ¬μ©νλ ν΄λ ꡬ쑰 μμμ
λλ€.
μ€μ νμΌλͺ
κ³Ό νμμ μλ³Έμ νμΈν ν λ°μν©λλ€.
```text
truthfulness-benchmark/
βββ README.md
βββ general-knowledge/
β βββ test.json
βββ numerical/
β βββ test.json
βββ fact-check/
β βββ test.json
βββ multimodal/
βββ test.json
βββ images/
```
ν΄λΉ ν΄λμ README μλ¨μ μλ μ€μ μ λ£μΌλ©΄ λ€ νλͺ©μ λ°λ‘ μ νν μ μμ΅λλ€.
κ²½λ‘λ μ μμμ ν΄λΉνλ©°, μ€μ νμΌ κ²½λ‘μ μΌμΉμμΌμΌ ν©λλ€.
```yaml
---
language:
- ko
configs:
- config_name: general_knowledge
data_files:
- split: test
path: general-knowledge/test.json
- config_name: numerical
data_files:
- split: test
path: numerical/test.json
- config_name: fact_check
data_files:
- split: test
path: fact-check/test.json
- config_name: multimodal
data_files:
- split: test
path: multimodal/test.json
---
```
## μλ²λ¦° Benchmark
```text
sovereign-benchmark/
βββ README.md
βββ test.json
```
μλ³Έμ΄ μ¬λ¬ νμΌ λλ νμ νκ° νλͺ©μΌλ‘ λλμ΄ μλ€λ©΄ κ·Έ ꡬ쑰λ₯Ό μ μ§νκ³ ,
νμμ λ°λΌ μλ²λ¦° μ μ₯μ μμλ configurationsλ₯Ό μ μν©λλ€.
## μ
λ‘λ
1. <https://huggingface.co/new-dataset>μμ κ° Dataset μ μ₯μλ₯Ό μμ±ν©λλ€.
2. 곡κ°ν μλ³Έ λ°μ΄ν°μ λ°μ΄ν° μ€λͺ
READMEλ₯Ό κ°κ°μ λ°°ν¬ ν΄λμ μ€λΉν©λλ€.
READMEμλ νλͺ©λ³ μΆμ², νλ, νκ° λ°©μ, λΌμ΄μ μ€λ₯Ό κΈ°μ¬ν©λλ€.
3. λ€μ λͺ
λ Ήμ `/μ€μ /κ²½λ‘/` λΆλΆμ μ€λΉν ν΄λ κ²½λ‘λ‘ λ°κΏ μ€νν©λλ€.
```bash
hf upload DoolyKim22/Truthfulness-Benchmark \
"/μ€μ /κ²½λ‘/truthfulness-benchmark" . --repo-type dataset
hf upload DoolyKim22/Sovereign-Benchmark \
"/μ€μ /κ²½λ‘/sovereign-benchmark" . --repo-type dataset
```
μ΄λ―Έμ§λ annotationμμ μ°Έμ‘°νλ μλ κ²½λ‘λ₯Ό μ μ§νμ¬ ν¨κ» μ
λ‘λν©λλ€.
`static-space/front/data/benchmark.json`μ λͺ¨λΈλ³ μ μ νμΌμ΄λ©° νκ° μλ³Έμ΄ μλλλ€.
νμ¬ μ΄ μμ
ν΄λμμ λ€μ― νλͺ©μ μλ³Έ μμΉλ νμΈλμ§ μμμ΅λλ€.
μλ³Έ κ²½λ‘λ₯Ό νμΈν λ€ λ°°ν¬ ν΄λλ₯Ό ꡬμ±νκ³ μ΄λ―Έμ§ μ°Έμ‘°μ λ°μ΄ν° νμμ κ²μ¦ν΄μΌ ν©λλ€.
μ°Έκ³ : [λ°μ΄ν°μ
μ
λ‘λ](https://huggingface.co/docs/hub/datasets-adding),
[μ¬λ¬ λ°μ΄ν°μ
ꡬμ±](https://huggingface.co/docs/hub/datasets-data-files-configuration).
|