luke0709 commited on
Commit
68f5f5e
·
1 Parent(s): ac7a77a

add git address

Browse files
Files changed (4) hide show
  1. DATASETS.md +95 -0
  2. README.md +1 -85
  3. index.html +6 -1
  4. upload_space.py +53 -0
DATASETS.md ADDED
@@ -0,0 +1,95 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Benchmark 데이터셋 배포 구성
2
+
3
+ Hugging Face Dataset 저장소를 아래 단위로 관리합니다. 개인 계정에서 먼저 배포하고
4
+ 검증한 뒤, 필요하면 각 저장소를 `KETI-NLP` 조직으로 이전합니다.
5
+ 아래 저장소 ID와 경로는 배포할 때 사용할 제안이며, 아직 생성된 저장소를 의미하지 않습니다.
6
+
7
+ | 표시 이름 | 개인 저장소 ID 제안 | 포함 데이터셋 |
8
+ | --- | --- | --- |
9
+ | 진실성 Benchmark | `DoolyKim22/Truthfulness-Benchmark` | 일반 상식, 수치 정보, 팩트 체크, 멀티 모달 |
10
+ | 소버린 Benchmark | `DoolyKim22/Sovereign-Benchmark` | 소버린 |
11
+ | K-Prism | `DoolyKim22/K-Prism` | K-Prism Text / Image 트랙 |
12
+
13
+ 진실성의 네 항목은 하나의 저장소 안에서 별도의 configuration으로 제공합니다.
14
+ 각 항목의 원본 필드와 정답 체계는 유지하며, 하나의 JSON으로 합치지 않습니다.
15
+ 이 구성은 데이터 배포 단위이며, 리더보드의 평가 항목이나 평균 계산은 변경하지 않습니다.
16
+
17
+ ## 진실성 Benchmark
18
+
19
+ 다음은 JSON 형식의 평가 데이터를 사용하는 폴더 구조 예시입니다.
20
+ 실제 파일명과 형식은 원본을 확인한 후 반영합니다.
21
+
22
+ ```text
23
+ truthfulness-benchmark/
24
+ ├── README.md
25
+ ├── general-knowledge/
26
+ │ └── test.json
27
+ ├── numerical/
28
+ │ └── test.json
29
+ ├── fact-check/
30
+ │ └── test.json
31
+ └── multimodal/
32
+ ├── test.json
33
+ └── images/
34
+ ```
35
+
36
+ 해당 폴더의 README 상단에 아래 설정을 넣으면 네 항목을 따로 선택할 수 있습니다.
37
+ 경로는 위 예시에 해당하며, 실제 파일 경로와 일치시켜야 합니다.
38
+
39
+ ```yaml
40
+ ---
41
+ language:
42
+ - ko
43
+ configs:
44
+ - config_name: general_knowledge
45
+ data_files:
46
+ - split: test
47
+ path: general-knowledge/test.json
48
+ - config_name: numerical
49
+ data_files:
50
+ - split: test
51
+ path: numerical/test.json
52
+ - config_name: fact_check
53
+ data_files:
54
+ - split: test
55
+ path: fact-check/test.json
56
+ - config_name: multimodal
57
+ data_files:
58
+ - split: test
59
+ path: multimodal/test.json
60
+ ---
61
+ ```
62
+
63
+ ## 소버린 Benchmark
64
+
65
+ ```text
66
+ sovereign-benchmark/
67
+ ├── README.md
68
+ └── test.json
69
+ ```
70
+
71
+ 원본이 여러 파일 또는 하위 평가 항목으로 나뉘어 있다면 그 구조를 유지하고,
72
+ 필요에 따라 소버린 저장소 안에도 configurations를 정의합니다.
73
+
74
+ ## 업로드
75
+
76
+ 1. <https://huggingface.co/new-dataset>에서 각 Dataset 저장소를 생성합니다.
77
+ 2. 공개할 원본 데이터와 데이터 설명 README를 각각의 배포 폴더에 준비합니다.
78
+ README에는 항목별 출처, 필드, 평가 방식, 라이선스를 기재합니다.
79
+ 3. 다음 명령의 `/실제/경로/` 부분을 준비한 폴더 경로로 바꿔 실행합니다.
80
+
81
+ ```bash
82
+ hf upload DoolyKim22/Truthfulness-Benchmark \
83
+ "/실제/경로/truthfulness-benchmark" . --repo-type dataset
84
+
85
+ hf upload DoolyKim22/Sovereign-Benchmark \
86
+ "/실제/경로/sovereign-benchmark" . --repo-type dataset
87
+ ```
88
+
89
+ 이미지는 annotation에서 참조하는 상대 경로를 유지하여 함께 업로드합니다.
90
+ `static-space/front/data/benchmark.json`은 모델별 점수 파일이며 평가 원본이 아닙니다.
91
+ 현재 이 작업 폴더에서 다섯 항목의 원본 위치는 확인되지 않았습니다.
92
+ 원본 경로를 확인한 뒤 배포 폴더를 구성하고 이미지 참조와 데이터 형식을 검증해야 합니다.
93
+
94
+ 참고: [데이터셋 업로드](https://huggingface.co/docs/hub/datasets-adding),
95
+ [여러 데이터셋 구성](https://huggingface.co/docs/hub/datasets-data-files-configuration).
README.md CHANGED
@@ -6,88 +6,4 @@ colorTo: gray
6
  sdk: static
7
  app_file: index.html
8
  pinned: false
9
- ---
10
-
11
- # KETI Leaderboard
12
-
13
- A static leaderboard for published KETI model evaluations. The browser loads
14
- `front/data/benchmark.json` directly; no Python server, Docker runtime, access token,
15
- or build step is needed. Dataset selection, model search, provider filtering,
16
- comparison charts, theme selection, and CSV export are available.
17
-
18
- The initial snapshot contains the 19 models and the 9 displayed datasets retrieved
19
- from the existing public leaderboard during migration. Scores and evaluation
20
- timestamps are preserved. The snapshot records its source and capture time;
21
- internal evaluation job identifiers and error messages are not included.
22
-
23
- ## Run locally
24
-
25
- From the repository root:
26
-
27
- ```bash
28
- python3 -m http.server 7860 --bind 127.0.0.1
29
- ```
30
-
31
- Open <http://127.0.0.1:7860/>. Use an HTTP server rather than opening `index.html`
32
- as a local file, because the browser loads JavaScript modules and JSON.
33
- Tailwind, Chart.js, and fonts currently load from their existing public CDNs.
34
-
35
- ## Verify the static page
36
-
37
- With Playwright and Chrome/Chromium installed, run:
38
-
39
- ```bash
40
- python3 -m unittest discover -s tests -v
41
- ```
42
-
43
- Set `CHROME_PATH` if the browser is not on `PATH`. These checks use a temporary
44
- local HTTP server and cover published scores and means, filters, comparison charts,
45
- CSV export, mobile controls, preferences, loading failures, and malformed data.
46
- Internet access is needed for the existing CDN assets in the chart test.
47
-
48
- ## Update published results
49
-
50
- 1. Evaluate models outside this Space.
51
- 2. Update `front/data/benchmark.json` with the resulting scores.
52
- 3. Commit and push the changed data file to the Space. Visitors can use **Refresh**
53
- to load the published version again.
54
-
55
- The data file accepts either the legacy `benchmark.json` array of models or an
56
- object with `models` and an optional `datasets` array. Each model has `provider`,
57
- `name`, `repo`, `is_multimodal`, `updated_at`, and `scores`. Each score has
58
- `dataset_name`, numeric `score`, and `metric_type` (`raw` or `llm-as-judge`).
59
- The `datasets` array sets the displayed columns and their order; without it,
60
- the canonical datasets and any additional scored datasets are shown.
61
-
62
- The mean calculation retains the previous behavior: all selected, supported
63
- datasets must have scores. Unsupported multimodal datasets are excluded for
64
- text-only models. Missing means are displayed as `—`. A failed data request shows
65
- an error; it never substitutes generated scores. A failed refresh keeps the last
66
- successfully loaded results on screen.
67
-
68
- For an independently updated, **public** JSON source, change `dataUrl` in
69
- `front/config.mjs` to its HTTPS URL. That server must allow browser cross-origin
70
- requests (CORS). No automatic fallback to another source is applied. The default
71
- bundled snapshot remains independent of the personal Hugging Face account.
72
- All browser configuration and bundled files are public: do not put tokens or
73
- private evaluation data in them.
74
-
75
- ## Deploy as a new Hugging Face Space
76
-
77
- This directory is an independent Git repository containing only the static app.
78
- The existing Docker Space at
79
- https://huggingface.co/spaces/DoolyKim22/KETI_Leaderboard remains separate.
80
-
81
- The deployment target is the new organization-owned Space
82
- [KETI-NLP/Benchmark_Leaderboard](https://huggingface.co/spaces/KETI-NLP/Benchmark_Leaderboard),
83
- using the **Static** SDK. The README metadata already specifies
84
- `sdk: static` and `app_file: index.html` as described in the
85
- [Static Spaces documentation](https://huggingface.co/docs/hub/spaces-sdks-static).
86
- Do not use the existing Docker Space as this repository's remote.
87
-
88
- The snapshot is independent of the running Docker app. Updates published to the
89
- old app after the snapshot was captured are not automatically copied here.
90
- Update `front/data/benchmark.json` and publish a new commit to refresh this app.
91
-
92
- Model submissions, automatic evaluations, and dataset backup workers are not part
93
- of the static site. Browser configuration lives in `front/config.mjs`.
 
6
  sdk: static
7
  app_file: index.html
8
  pinned: false
9
+ ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
index.html CHANGED
@@ -129,9 +129,10 @@
129
  <!-- Main -->
130
  <main class="mx-auto max-w-7xl px-4 md:px-6 py-6 md:py-8 space-y-6 md:space-y-8">
131
  <!-- Tabs -->
132
- <div class="flex gap-2 md:gap-3">
133
  <button class="tab tab-active" data-tab="leaderboard">🏆 Leaderboard</button>
134
  <button class="tab" data-tab="about">ℹ️ About</button>
 
135
  </div>
136
 
137
  <div id="dataStatus" class="card p-4 text-sm" role="status" aria-live="polite" data-state="loading">Loading published results…</div>
@@ -230,6 +231,10 @@
230
  <p class="text-sm text-slate-600 dark:text-slate-300">
231
  This leaderboard presents published model evaluations on KETI's ethicality and veracity benchmark datasets.
232
  </p>
 
 
 
 
233
  <p class="text-sm text-slate-600 dark:text-slate-300">
234
  Model submissions and automatic evaluations are not available on this page. Results are updated when the maintainers publish a new evaluation.
235
  </p>
 
129
  <!-- Main -->
130
  <main class="mx-auto max-w-7xl px-4 md:px-6 py-6 md:py-8 space-y-6 md:space-y-8">
131
  <!-- Tabs -->
132
+ <div class="flex flex-wrap gap-2 md:gap-3">
133
  <button class="tab tab-active" data-tab="leaderboard">🏆 Leaderboard</button>
134
  <button class="tab" data-tab="about">ℹ️ About</button>
135
+ <a class="btn btn-ghost" href="https://github.com/alsgur0720/K-Prism" target="_blank" rel="noopener noreferrer">↗ K-Prism Code</a>
136
  </div>
137
 
138
  <div id="dataStatus" class="card p-4 text-sm" role="status" aria-live="polite" data-state="loading">Loading published results…</div>
 
231
  <p class="text-sm text-slate-600 dark:text-slate-300">
232
  This leaderboard presents published model evaluations on KETI's ethicality and veracity benchmark datasets.
233
  </p>
234
+ <p class="text-sm text-slate-600 dark:text-slate-300">
235
+ Detailed K-Prism evaluation code and documentation:
236
+ <a class="underline" href="https://github.com/alsgur0720/K-Prism" target="_blank" rel="noopener noreferrer">alsgur0720/K-Prism on GitHub</a>.
237
+ </p>
238
  <p class="text-sm text-slate-600 dark:text-slate-300">
239
  Model submissions and automatic evaluations are not available on this page. Results are updated when the maintainers publish a new evaluation.
240
  </p>
upload_space.py ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Upload this static app to an existing Static Space using the saved HF login."""
2
+
3
+ import argparse
4
+ from pathlib import Path
5
+
6
+ from huggingface_hub import HfApi
7
+
8
+
9
+ ROOT = Path(__file__).resolve().parent
10
+ FILES = [
11
+ '.gitignore',
12
+ 'README.md',
13
+ 'index.html',
14
+ 'front/config.mjs',
15
+ 'front/data.mjs',
16
+ 'front/leaderboard.mjs',
17
+ 'front/data/benchmark.json',
18
+ 'tests/test_static_space.py',
19
+ 'upload_space.py',
20
+ ]
21
+
22
+
23
+ def main():
24
+ parser = argparse.ArgumentParser(description=__doc__)
25
+ parser.add_argument('--repo-id', default='DoolyKim22/Benchmark_Leaderboard')
26
+ parser.add_argument('--message', default='Update static benchmark leaderboard')
27
+ args = parser.parse_args()
28
+
29
+ for name in FILES:
30
+ if not (ROOT / name).is_file():
31
+ parser.error('Missing deployment file: ' + name)
32
+
33
+ api = HfApi()
34
+ space = api.space_info(args.repo_id)
35
+ if space.sdk != 'static':
36
+ parser.error('Upload target must already be a Static Space: ' + args.repo_id)
37
+
38
+ # upload_folder writes to the existing repository without create_repo().
39
+ # Older hf upload versions send space_sdk="gradio" in that extra request.
40
+ commit = api.upload_folder(
41
+ repo_id=args.repo_id,
42
+ repo_type='space',
43
+ folder_path=ROOT,
44
+ allow_patterns=FILES,
45
+ parent_commit=space.sha,
46
+ commit_message=args.message,
47
+ )
48
+ print('Uploaded:', commit.commit_url)
49
+ print('Space: https://huggingface.co/spaces/' + args.repo_id)
50
+
51
+
52
+ if __name__ == '__main__':
53
+ main()