Nanthasit commited on
Commit
7cfc347
·
verified ·
1 Parent(s): a373835

health-eval: automated health check 2026-07-30

Browse files
Files changed (1) hide show
  1. .eval_results/health-check.yaml +105 -0
.eval_results/health-check.yaml ADDED
@@ -0,0 +1,105 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Health Check Report
2
+ # Generated: 2026-07-30T23:06 UTC
3
+ # Model: Nanthasit/sakthai-coder-1.5b
4
+ # Tool: sakthai-agent health-eval (cron)
5
+
6
+ model: Nanthasit/sakthai-coder-1.5b
7
+ eval_timestamp: 2026-07-30T23:06:00Z
8
+
9
+ metrics:
10
+ popularity:
11
+ downloads: 93
12
+ likes: 0
13
+ last_modified: 2026-07-30T22:56:24.000Z
14
+ days_since_last_update: 0
15
+ metadata:
16
+ pipeline_tag: text-generation
17
+ library_name: transformers
18
+ license: apache-2.0
19
+ base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
20
+ private: false
21
+ tags:
22
+ - text-generation
23
+ - transformers
24
+ - qwen
25
+ - coder
26
+ - sakthai
27
+ - finetuned
28
+ - tool-calling
29
+ - function-calling
30
+ - 1.5b
31
+ - llama.cpp
32
+ datasets:
33
+ - Nanthasit/sakthai-combined-v6
34
+ - Nanthasit/sakthai-combined-v7
35
+ - Nanthasit/sakthai-irrelevance-supplement
36
+ datasets_count: 3
37
+
38
+ files:
39
+ total_siblings: 1563
40
+ model_weight_files: 1
41
+ model_weights:
42
+ - file: qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
43
+ size_bytes: 1117320768
44
+ size_gb: 1.041
45
+ total_model_size_gb: 1.041
46
+ contamination: true
47
+ contamination_note: "1563 siblings detected. Repo contains a full .venv/ directory (1254 .py files, pip packages, .dist-info). Only 1 actual model weight file. Recommends cleanup: remove .venv/ from repo."
48
+
49
+ benchmarks:
50
+ model_index_present: true
51
+ results:
52
+ - dataset: HumanEval
53
+ task: text-generation
54
+ metric: pass@1 (base model reference)
55
+ value: 74.4
56
+ verified: false
57
+ - dataset: MBPP
58
+ task: text-generation
59
+ metric: pass@1 (base model reference)
60
+ value: 71.2
61
+ verified: false
62
+ - dataset: MultiPL-E (Python)
63
+ task: text-generation
64
+ metric: pass@1 (base model reference)
65
+ value: 65.3
66
+ verified: false
67
+ - dataset: SakThai Coding Suite (internal)
68
+ task: text-generation
69
+ metric: pass@1 (fine-tuned model, internal)
70
+ value: 100.0
71
+ verified: false
72
+ source: internal-local-llama-cpp-2026-07-25
73
+ benchmark_quality:
74
+ - "Internal benchmark (SakThai Coding Suite) shows 100% pass@1 but unverified — likely overfit or too narrow"
75
+ - "Base model reference scores from Qwen2.5-Coder-1.5B-Instruct used as comparison; no verified fine-tuned scores"
76
+ - "No multi-trial methodology reported; single-trial results may be unreliable"
77
+
78
+ health_score:
79
+ overall: 6.5
80
+ max: 10
81
+ factors:
82
+ popularity_downloads: 2
83
+ likes: 0
84
+ metadata_quality: 8
85
+ model_index: 7
86
+ file_hygiene: 3
87
+ recency: 9
88
+ concerns:
89
+ - "Repo severely contaminated with .venv/ directory (1254 transient .py files)"
90
+ - "0 likes — no community engagement yet"
91
+ - "Benchmark scores unverified (marked verified: false)"
92
+ - "Only 1 model weight format (GGUF Q4_K_M); no safetensors variant"
93
+ recommendations:
94
+ - action: REMOVE_VENV
95
+ priority: high
96
+ detail: "Delete .venv/ from repo — it adds 1254+ useless files and obscures real model content"
97
+ - action: ADD_SAFETENSORS
98
+ priority: medium
99
+ detail: "Consider adding a safetensors version for Transformers-native loading"
100
+ - action: VERIFY_BENCHMARKS
101
+ priority: medium
102
+ detail: "Run multi-trial benchmarks and mark as verified"
103
+ - action: ADD_MODEL_CARD_SECTIONS
104
+ priority: low
105
+ detail: "Add usage example, recommended prompt format, and hardware requirements"