YAML Metadata Error:Invalid Eval Result format in .eval_results/health-check.yaml

Check out the documentation for more information.

Show details
✖ Invalid input: expected array, received object
sakthai-coder-1.5b / .eval_results /health-check.yaml
Nanthasit's picture
health-eval: automated health check 2026-07-30
7cfc347 verified
Raw History Blame
3.34 kB
# Health Check Report
# Generated: 2026-07-30T23:06 UTC
# Model: Nanthasit/sakthai-coder-1.5b
# Tool: sakthai-agent health-eval (cron)
model: Nanthasit/sakthai-coder-1.5b
eval_timestamp: 2026-07-30T23:06:00Z
metrics:
popularity:
downloads: 93
likes: 0
last_modified: 2026-07-30T22:56:24.000Z
days_since_last_update: 0
metadata:
pipeline_tag: text-generation
library_name: transformers
license: apache-2.0
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
private: false
tags:
- text-generation
- transformers
- qwen
- coder
- sakthai
- finetuned
- tool-calling
- function-calling
- 1.5b
- llama.cpp
datasets:
- Nanthasit/sakthai-combined-v6
- Nanthasit/sakthai-combined-v7
- Nanthasit/sakthai-irrelevance-supplement
datasets_count: 3
files:
total_siblings: 1563
model_weight_files: 1
model_weights:
- file: qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
size_bytes: 1117320768
size_gb: 1.041
total_model_size_gb: 1.041
contamination: true
contamination_note: "1563 siblings detected. Repo contains a full .venv/ directory (1254 .py files, pip packages, .dist-info). Only 1 actual model weight file. Recommends cleanup: remove .venv/ from repo."
benchmarks:
model_index_present: true
results:
- dataset: HumanEval
task: text-generation
metric: pass@1 (base model reference)
value: 74.4
verified: false
- dataset: MBPP
task: text-generation
metric: pass@1 (base model reference)
value: 71.2
verified: false
- dataset: MultiPL-E (Python)
task: text-generation
metric: pass@1 (base model reference)
value: 65.3
verified: false
- dataset: SakThai Coding Suite (internal)
task: text-generation
metric: pass@1 (fine-tuned model, internal)
value: 100.0
verified: false
source: internal-local-llama-cpp-2026-07-25
benchmark_quality:
- "Internal benchmark (SakThai Coding Suite) shows 100% pass@1 but unverified — likely overfit or too narrow"
- "Base model reference scores from Qwen2.5-Coder-1.5B-Instruct used as comparison; no verified fine-tuned scores"
- "No multi-trial methodology reported; single-trial results may be unreliable"
health_score:
overall: 6.5
max: 10
factors:
popularity_downloads: 2
likes: 0
metadata_quality: 8
model_index: 7
file_hygiene: 3
recency: 9
concerns:
- "Repo severely contaminated with .venv/ directory (1254 transient .py files)"
- "0 likes — no community engagement yet"
- "Benchmark scores unverified (marked verified: false)"
- "Only 1 model weight format (GGUF Q4_K_M); no safetensors variant"
recommendations:
- action: REMOVE_VENV
priority: high
detail: "Delete .venv/ from repo — it adds 1254+ useless files and obscures real model content"
- action: ADD_SAFETENSORS
priority: medium
detail: "Consider adding a safetensors version for Transformers-native loading"
- action: VERIFY_BENCHMARKS
priority: medium
detail: "Run multi-trial benchmarks and mark as verified"
- action: ADD_MODEL_CARD_SECTIONS
priority: low
detail: "Add usage example, recommended prompt format, and hardware requirements"