Spaces:
Running
Running
luke0709 commited on
Commit ·
68f5f5e
1
Parent(s): ac7a77a
add git address
Browse files- DATASETS.md +95 -0
- README.md +1 -85
- index.html +6 -1
- upload_space.py +53 -0
DATASETS.md
ADDED
|
@@ -0,0 +1,95 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Benchmark 데이터셋 배포 구성
|
| 2 |
+
|
| 3 |
+
Hugging Face Dataset 저장소를 아래 단위로 관리합니다. 개인 계정에서 먼저 배포하고
|
| 4 |
+
검증한 뒤, 필요하면 각 저장소를 `KETI-NLP` 조직으로 이전합니다.
|
| 5 |
+
아래 저장소 ID와 경로는 배포할 때 사용할 제안이며, 아직 생성된 저장소를 의미하지 않습니다.
|
| 6 |
+
|
| 7 |
+
| 표시 이름 | 개인 저장소 ID 제안 | 포함 데이터셋 |
|
| 8 |
+
| --- | --- | --- |
|
| 9 |
+
| 진실성 Benchmark | `DoolyKim22/Truthfulness-Benchmark` | 일반 상식, 수치 정보, 팩트 체크, 멀티 모달 |
|
| 10 |
+
| 소버린 Benchmark | `DoolyKim22/Sovereign-Benchmark` | 소버린 |
|
| 11 |
+
| K-Prism | `DoolyKim22/K-Prism` | K-Prism Text / Image 트랙 |
|
| 12 |
+
|
| 13 |
+
진실성의 네 항목은 하나의 저장소 안에서 별도의 configuration으로 제공합니다.
|
| 14 |
+
각 항목의 원본 필드와 정답 체계는 유지하며, 하나의 JSON으로 합치지 않습니다.
|
| 15 |
+
이 구성은 데이터 배포 단위이며, 리더보드의 평가 항목이나 평균 계산은 변경하지 않습니다.
|
| 16 |
+
|
| 17 |
+
## 진실성 Benchmark
|
| 18 |
+
|
| 19 |
+
다음은 JSON 형식의 평가 데이터를 사용하는 폴더 구조 예시입니다.
|
| 20 |
+
실제 파일명과 형식은 원본을 확인한 후 반영합니다.
|
| 21 |
+
|
| 22 |
+
```text
|
| 23 |
+
truthfulness-benchmark/
|
| 24 |
+
├── README.md
|
| 25 |
+
├── general-knowledge/
|
| 26 |
+
│ └── test.json
|
| 27 |
+
├── numerical/
|
| 28 |
+
│ └── test.json
|
| 29 |
+
├── fact-check/
|
| 30 |
+
│ └── test.json
|
| 31 |
+
└── multimodal/
|
| 32 |
+
├── test.json
|
| 33 |
+
└── images/
|
| 34 |
+
```
|
| 35 |
+
|
| 36 |
+
해당 폴더의 README 상단에 아래 설정을 넣으면 네 항목을 따로 선택할 수 있습니다.
|
| 37 |
+
경로는 위 예시에 해당하며, 실제 파일 경로와 일치시켜야 합니다.
|
| 38 |
+
|
| 39 |
+
```yaml
|
| 40 |
+
---
|
| 41 |
+
language:
|
| 42 |
+
- ko
|
| 43 |
+
configs:
|
| 44 |
+
- config_name: general_knowledge
|
| 45 |
+
data_files:
|
| 46 |
+
- split: test
|
| 47 |
+
path: general-knowledge/test.json
|
| 48 |
+
- config_name: numerical
|
| 49 |
+
data_files:
|
| 50 |
+
- split: test
|
| 51 |
+
path: numerical/test.json
|
| 52 |
+
- config_name: fact_check
|
| 53 |
+
data_files:
|
| 54 |
+
- split: test
|
| 55 |
+
path: fact-check/test.json
|
| 56 |
+
- config_name: multimodal
|
| 57 |
+
data_files:
|
| 58 |
+
- split: test
|
| 59 |
+
path: multimodal/test.json
|
| 60 |
+
---
|
| 61 |
+
```
|
| 62 |
+
|
| 63 |
+
## 소버린 Benchmark
|
| 64 |
+
|
| 65 |
+
```text
|
| 66 |
+
sovereign-benchmark/
|
| 67 |
+
├── README.md
|
| 68 |
+
└── test.json
|
| 69 |
+
```
|
| 70 |
+
|
| 71 |
+
원본이 여러 파일 또는 하위 평가 항목으로 나뉘어 있다면 그 구조를 유지하고,
|
| 72 |
+
필요에 따라 소버린 저장소 안에도 configurations를 정의합니다.
|
| 73 |
+
|
| 74 |
+
## 업로드
|
| 75 |
+
|
| 76 |
+
1. <https://huggingface.co/new-dataset>에서 각 Dataset 저장소를 생성합니다.
|
| 77 |
+
2. 공개할 원본 데이터와 데이터 설명 README를 각각의 배포 폴더에 준비합니다.
|
| 78 |
+
README에는 항목별 출처, 필드, 평가 방식, 라이선스를 기재합니다.
|
| 79 |
+
3. 다음 명령의 `/실제/경로/` 부분을 준비한 폴더 경로로 바꿔 실행합니다.
|
| 80 |
+
|
| 81 |
+
```bash
|
| 82 |
+
hf upload DoolyKim22/Truthfulness-Benchmark \
|
| 83 |
+
"/실제/경로/truthfulness-benchmark" . --repo-type dataset
|
| 84 |
+
|
| 85 |
+
hf upload DoolyKim22/Sovereign-Benchmark \
|
| 86 |
+
"/실제/경로/sovereign-benchmark" . --repo-type dataset
|
| 87 |
+
```
|
| 88 |
+
|
| 89 |
+
이미지는 annotation에서 참조하는 상대 경로를 유지하여 함께 업로드합니다.
|
| 90 |
+
`static-space/front/data/benchmark.json`은 모델별 점수 파일이며 평가 원본이 아닙니다.
|
| 91 |
+
현재 이 작업 폴더에서 다섯 항목의 원본 위치는 확인되지 않았습니다.
|
| 92 |
+
원본 경로를 확인한 뒤 배포 폴더를 구성하고 이미지 참조와 데이터 형식을 검증해야 합니다.
|
| 93 |
+
|
| 94 |
+
참고: [데이터셋 업로드](https://huggingface.co/docs/hub/datasets-adding),
|
| 95 |
+
[여러 데이터셋 구성](https://huggingface.co/docs/hub/datasets-data-files-configuration).
|
README.md
CHANGED
|
@@ -6,88 +6,4 @@ colorTo: gray
|
|
| 6 |
sdk: static
|
| 7 |
app_file: index.html
|
| 8 |
pinned: false
|
| 9 |
-
---
|
| 10 |
-
|
| 11 |
-
# KETI Leaderboard
|
| 12 |
-
|
| 13 |
-
A static leaderboard for published KETI model evaluations. The browser loads
|
| 14 |
-
`front/data/benchmark.json` directly; no Python server, Docker runtime, access token,
|
| 15 |
-
or build step is needed. Dataset selection, model search, provider filtering,
|
| 16 |
-
comparison charts, theme selection, and CSV export are available.
|
| 17 |
-
|
| 18 |
-
The initial snapshot contains the 19 models and the 9 displayed datasets retrieved
|
| 19 |
-
from the existing public leaderboard during migration. Scores and evaluation
|
| 20 |
-
timestamps are preserved. The snapshot records its source and capture time;
|
| 21 |
-
internal evaluation job identifiers and error messages are not included.
|
| 22 |
-
|
| 23 |
-
## Run locally
|
| 24 |
-
|
| 25 |
-
From the repository root:
|
| 26 |
-
|
| 27 |
-
```bash
|
| 28 |
-
python3 -m http.server 7860 --bind 127.0.0.1
|
| 29 |
-
```
|
| 30 |
-
|
| 31 |
-
Open <http://127.0.0.1:7860/>. Use an HTTP server rather than opening `index.html`
|
| 32 |
-
as a local file, because the browser loads JavaScript modules and JSON.
|
| 33 |
-
Tailwind, Chart.js, and fonts currently load from their existing public CDNs.
|
| 34 |
-
|
| 35 |
-
## Verify the static page
|
| 36 |
-
|
| 37 |
-
With Playwright and Chrome/Chromium installed, run:
|
| 38 |
-
|
| 39 |
-
```bash
|
| 40 |
-
python3 -m unittest discover -s tests -v
|
| 41 |
-
```
|
| 42 |
-
|
| 43 |
-
Set `CHROME_PATH` if the browser is not on `PATH`. These checks use a temporary
|
| 44 |
-
local HTTP server and cover published scores and means, filters, comparison charts,
|
| 45 |
-
CSV export, mobile controls, preferences, loading failures, and malformed data.
|
| 46 |
-
Internet access is needed for the existing CDN assets in the chart test.
|
| 47 |
-
|
| 48 |
-
## Update published results
|
| 49 |
-
|
| 50 |
-
1. Evaluate models outside this Space.
|
| 51 |
-
2. Update `front/data/benchmark.json` with the resulting scores.
|
| 52 |
-
3. Commit and push the changed data file to the Space. Visitors can use **Refresh**
|
| 53 |
-
to load the published version again.
|
| 54 |
-
|
| 55 |
-
The data file accepts either the legacy `benchmark.json` array of models or an
|
| 56 |
-
object with `models` and an optional `datasets` array. Each model has `provider`,
|
| 57 |
-
`name`, `repo`, `is_multimodal`, `updated_at`, and `scores`. Each score has
|
| 58 |
-
`dataset_name`, numeric `score`, and `metric_type` (`raw` or `llm-as-judge`).
|
| 59 |
-
The `datasets` array sets the displayed columns and their order; without it,
|
| 60 |
-
the canonical datasets and any additional scored datasets are shown.
|
| 61 |
-
|
| 62 |
-
The mean calculation retains the previous behavior: all selected, supported
|
| 63 |
-
datasets must have scores. Unsupported multimodal datasets are excluded for
|
| 64 |
-
text-only models. Missing means are displayed as `—`. A failed data request shows
|
| 65 |
-
an error; it never substitutes generated scores. A failed refresh keeps the last
|
| 66 |
-
successfully loaded results on screen.
|
| 67 |
-
|
| 68 |
-
For an independently updated, **public** JSON source, change `dataUrl` in
|
| 69 |
-
`front/config.mjs` to its HTTPS URL. That server must allow browser cross-origin
|
| 70 |
-
requests (CORS). No automatic fallback to another source is applied. The default
|
| 71 |
-
bundled snapshot remains independent of the personal Hugging Face account.
|
| 72 |
-
All browser configuration and bundled files are public: do not put tokens or
|
| 73 |
-
private evaluation data in them.
|
| 74 |
-
|
| 75 |
-
## Deploy as a new Hugging Face Space
|
| 76 |
-
|
| 77 |
-
This directory is an independent Git repository containing only the static app.
|
| 78 |
-
The existing Docker Space at
|
| 79 |
-
https://huggingface.co/spaces/DoolyKim22/KETI_Leaderboard remains separate.
|
| 80 |
-
|
| 81 |
-
The deployment target is the new organization-owned Space
|
| 82 |
-
[KETI-NLP/Benchmark_Leaderboard](https://huggingface.co/spaces/KETI-NLP/Benchmark_Leaderboard),
|
| 83 |
-
using the **Static** SDK. The README metadata already specifies
|
| 84 |
-
`sdk: static` and `app_file: index.html` as described in the
|
| 85 |
-
[Static Spaces documentation](https://huggingface.co/docs/hub/spaces-sdks-static).
|
| 86 |
-
Do not use the existing Docker Space as this repository's remote.
|
| 87 |
-
|
| 88 |
-
The snapshot is independent of the running Docker app. Updates published to the
|
| 89 |
-
old app after the snapshot was captured are not automatically copied here.
|
| 90 |
-
Update `front/data/benchmark.json` and publish a new commit to refresh this app.
|
| 91 |
-
|
| 92 |
-
Model submissions, automatic evaluations, and dataset backup workers are not part
|
| 93 |
-
of the static site. Browser configuration lives in `front/config.mjs`.
|
|
|
|
| 6 |
sdk: static
|
| 7 |
app_file: index.html
|
| 8 |
pinned: false
|
| 9 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
index.html
CHANGED
|
@@ -129,9 +129,10 @@
|
|
| 129 |
<!-- Main -->
|
| 130 |
<main class="mx-auto max-w-7xl px-4 md:px-6 py-6 md:py-8 space-y-6 md:space-y-8">
|
| 131 |
<!-- Tabs -->
|
| 132 |
-
<div class="flex gap-2 md:gap-3">
|
| 133 |
<button class="tab tab-active" data-tab="leaderboard">🏆 Leaderboard</button>
|
| 134 |
<button class="tab" data-tab="about">ℹ️ About</button>
|
|
|
|
| 135 |
</div>
|
| 136 |
|
| 137 |
<div id="dataStatus" class="card p-4 text-sm" role="status" aria-live="polite" data-state="loading">Loading published results…</div>
|
|
@@ -230,6 +231,10 @@
|
|
| 230 |
<p class="text-sm text-slate-600 dark:text-slate-300">
|
| 231 |
This leaderboard presents published model evaluations on KETI's ethicality and veracity benchmark datasets.
|
| 232 |
</p>
|
|
|
|
|
|
|
|
|
|
|
|
|
| 233 |
<p class="text-sm text-slate-600 dark:text-slate-300">
|
| 234 |
Model submissions and automatic evaluations are not available on this page. Results are updated when the maintainers publish a new evaluation.
|
| 235 |
</p>
|
|
|
|
| 129 |
<!-- Main -->
|
| 130 |
<main class="mx-auto max-w-7xl px-4 md:px-6 py-6 md:py-8 space-y-6 md:space-y-8">
|
| 131 |
<!-- Tabs -->
|
| 132 |
+
<div class="flex flex-wrap gap-2 md:gap-3">
|
| 133 |
<button class="tab tab-active" data-tab="leaderboard">🏆 Leaderboard</button>
|
| 134 |
<button class="tab" data-tab="about">ℹ️ About</button>
|
| 135 |
+
<a class="btn btn-ghost" href="https://github.com/alsgur0720/K-Prism" target="_blank" rel="noopener noreferrer">↗ K-Prism Code</a>
|
| 136 |
</div>
|
| 137 |
|
| 138 |
<div id="dataStatus" class="card p-4 text-sm" role="status" aria-live="polite" data-state="loading">Loading published results…</div>
|
|
|
|
| 231 |
<p class="text-sm text-slate-600 dark:text-slate-300">
|
| 232 |
This leaderboard presents published model evaluations on KETI's ethicality and veracity benchmark datasets.
|
| 233 |
</p>
|
| 234 |
+
<p class="text-sm text-slate-600 dark:text-slate-300">
|
| 235 |
+
Detailed K-Prism evaluation code and documentation:
|
| 236 |
+
<a class="underline" href="https://github.com/alsgur0720/K-Prism" target="_blank" rel="noopener noreferrer">alsgur0720/K-Prism on GitHub</a>.
|
| 237 |
+
</p>
|
| 238 |
<p class="text-sm text-slate-600 dark:text-slate-300">
|
| 239 |
Model submissions and automatic evaluations are not available on this page. Results are updated when the maintainers publish a new evaluation.
|
| 240 |
</p>
|
upload_space.py
ADDED
|
@@ -0,0 +1,53 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Upload this static app to an existing Static Space using the saved HF login."""
|
| 2 |
+
|
| 3 |
+
import argparse
|
| 4 |
+
from pathlib import Path
|
| 5 |
+
|
| 6 |
+
from huggingface_hub import HfApi
|
| 7 |
+
|
| 8 |
+
|
| 9 |
+
ROOT = Path(__file__).resolve().parent
|
| 10 |
+
FILES = [
|
| 11 |
+
'.gitignore',
|
| 12 |
+
'README.md',
|
| 13 |
+
'index.html',
|
| 14 |
+
'front/config.mjs',
|
| 15 |
+
'front/data.mjs',
|
| 16 |
+
'front/leaderboard.mjs',
|
| 17 |
+
'front/data/benchmark.json',
|
| 18 |
+
'tests/test_static_space.py',
|
| 19 |
+
'upload_space.py',
|
| 20 |
+
]
|
| 21 |
+
|
| 22 |
+
|
| 23 |
+
def main():
|
| 24 |
+
parser = argparse.ArgumentParser(description=__doc__)
|
| 25 |
+
parser.add_argument('--repo-id', default='DoolyKim22/Benchmark_Leaderboard')
|
| 26 |
+
parser.add_argument('--message', default='Update static benchmark leaderboard')
|
| 27 |
+
args = parser.parse_args()
|
| 28 |
+
|
| 29 |
+
for name in FILES:
|
| 30 |
+
if not (ROOT / name).is_file():
|
| 31 |
+
parser.error('Missing deployment file: ' + name)
|
| 32 |
+
|
| 33 |
+
api = HfApi()
|
| 34 |
+
space = api.space_info(args.repo_id)
|
| 35 |
+
if space.sdk != 'static':
|
| 36 |
+
parser.error('Upload target must already be a Static Space: ' + args.repo_id)
|
| 37 |
+
|
| 38 |
+
# upload_folder writes to the existing repository without create_repo().
|
| 39 |
+
# Older hf upload versions send space_sdk="gradio" in that extra request.
|
| 40 |
+
commit = api.upload_folder(
|
| 41 |
+
repo_id=args.repo_id,
|
| 42 |
+
repo_type='space',
|
| 43 |
+
folder_path=ROOT,
|
| 44 |
+
allow_patterns=FILES,
|
| 45 |
+
parent_commit=space.sha,
|
| 46 |
+
commit_message=args.message,
|
| 47 |
+
)
|
| 48 |
+
print('Uploaded:', commit.commit_url)
|
| 49 |
+
print('Space: https://huggingface.co/spaces/' + args.repo_id)
|
| 50 |
+
|
| 51 |
+
|
| 52 |
+
if __name__ == '__main__':
|
| 53 |
+
main()
|