How to use from
SGLang
Install from pip and serve model
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "palette-lab/songgot-m" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "palette-lab/songgot-m",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Use Docker images
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "palette-lab/songgot-m" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "palette-lab/songgot-m",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'
Quick Links

Songgot (์†ก๊ณณ)

A Korean-first tiny agentic model for tool calling on the device, trained from scratch by Hanish Keloth (Palette). Apache 2.0.

Try it on the device: https://hanishkeloth.github.io/songgot/app/ (runs in the browser, works offline after the first load).

Numbers

Kakao FunctionChat-Bench SingleCall (500 Korean items, 5 tool conditions), exact match on function name and arguments, scorer in the repo. Comparators run with identical tools and queries in their own documented formats.

model params exact 4_random 4_close 8_random 8_close all name only
Songgot-M (2 epochs, v8 set) 126M 45.0 43.0 26.0 34.0 18.0 33.2 73.8
Songgot (6B tokens, 3 epochs, v8 set) 50M 44.0 39.0 30.0 35.0 17.0 33.0 73.6
Songgot-nano (1 epoch) 39M 0.0 0.0 0.0 0.0 0.0 0.0 0.0
Needle 2 45M 0.0 0.0 0.0 0.0 0.0 0.0 0.0
FunctionGemma-270M 270M 3.0 5.0 1.0 1.0 1.0 2.2 36.2
Qwen3-0.6B 600M 48.0 49.0 45.0 37.0 37.0 43.2 70.8
Qwen3.5-0.8B 800M 51.0 48.0 41.0 52.0 34.0 45.2 73.6

Tokens per Hangul syllable on the same 100 queries: Songgot 0.90, Gemma 3 0.98, Qwen3 1.15, Needle 2 3.47.

Tokens per Hangul syllable

Call accuracy by condition

Pretraining loss

Status (2026-09-11 05:18)

Weights in this repo are Songgot-M, 2 epochs, v8 set: 16 layers, hidden 768, about 126M parameters, pretrained on 8xH100 (Modal) on 6B tokens, post-trained on the v8 set (five teacher-synthesis rounds, the last with confusable sibling tools), post-trained on the v2 tool-calling set. Call accuracy on FunctionChat-Bench SingleCall 33.2 percent (name only 73.8). GGUF exports (f16, Q8_0, Q4_K_M) are in this repo.

Format

<|system|>
[{"name": "set_alarm", "description": "์•Œ๋žŒ์„ ์„ค์ •ํ•ฉ๋‹ˆ๋‹ค.", "parameters": {...}}]
<|user|>
๋‚ด์ผ ์•„์นจ 7์‹œ์— ์•Œ๋žŒ ๋งž์ถฐ์ค˜
<|call|>
{"name":"set_alarm","arguments":{"time":"07:00"}}<|end|>

Tokenizer: SentencePiece BPE, 32k, byte fallback (tokenizer.model). Use sentencepiece directly; the special tokens live inside the vocabulary.

Data and provenance

fineweb-edu sample-10BT (ODC-By), Korean Wikipedia 20231101.ko (CC BY-SA 3.0; this model card carries the attribution and share-alike notice for that text), glaive-function-calling-v2 (Apache 2.0), template-generated Korean tool calls (Apache 2.0, in the repo). No closed-model outputs. FunctionChat-Bench was never used for training.

Limits

Single-call tool selection and argument extraction only. No multi-turn, no tool results, no free chat. Small models are finicky with rare tools and paraphrased values; validate every call in application code.

Downloads last month
547
Safetensors
Model size
0.1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Datasets used to train palette-lab/songgot-m