SebastianAldrin commited on
Commit
88d73ea
·
verified ·
1 Parent(s): 9ff9b5a

Link the Agent Society repo; point the dataset link at the current username

Browse files
Files changed (1) hide show
  1. README.md +4 -3
README.md CHANGED
@@ -22,7 +22,8 @@ runs locally and offline instead of calling a frontier model every tick.
22
  | Base | `Qwen/Qwen3.5-9B` |
23
  | Method | LoRA SFT — rank 16, alpha 32, dropout 0.05, all-linear, 3 epochs |
24
  | Teacher | Claude Sonnet |
25
- | Data | [Chanito91/agent-society-distill-v1](https://huggingface.co/datasets/Chanito91/agent-society-distill-v1) — 1,355 examples |
 
26
  | Format | GGUF, **Q4_K_M** (~5.4 GB) — runs in llama.cpp / Ollama |
27
 
28
  Given the agent's situation as a prompt (felt needs, a local block-map, bearings, the
@@ -39,7 +40,7 @@ llama-server -m Qwen3.5-9B-minecraft-distill-v1-Q4_K_M.gguf -c 8192 --jinja
39
 
40
  Send the decide prompt with `enable_thinking: false` and a JSON-schema response format
41
  (the model is trained to answer with the decision JSON only). Full system prompt and
42
- prompt format live in the Agent Society repository.
43
 
44
  On CUDA, add `--flash-attn off` — llama.cpp's `auto` default silently corrupts this
45
  hybrid-SSM architecture: the JSON shape survives but the words inside turn to noise.
@@ -52,7 +53,7 @@ hybrid-SSM architecture: the JSON shape survives but the words inside turn to no
52
  - The prompt isn't cacheable on this arch, so it re-reads the whole prompt every tick.
53
  Slow on a weak GPU or CPU.
54
  - Measured: on 250 held-out teacher decisions it picks the teacher's action **78.0%**
55
- of the time; the untuned base scores 56.8%. Method in the repo's evaluation doc.
56
 
57
  ## License
58
 
 
22
  | Base | `Qwen/Qwen3.5-9B` |
23
  | Method | LoRA SFT — rank 16, alpha 32, dropout 0.05, all-linear, 3 epochs |
24
  | Teacher | Claude Sonnet |
25
+ | Data | [SebastianAldrin/agent-society-distill-v1](https://huggingface.co/datasets/SebastianAldrin/agent-society-distill-v1) — 1,355 examples |
26
+ | Code | [Agent Society](https://github.com/sebastianaldrin/agent-society) |
27
  | Format | GGUF, **Q4_K_M** (~5.4 GB) — runs in llama.cpp / Ollama |
28
 
29
  Given the agent's situation as a prompt (felt needs, a local block-map, bearings, the
 
40
 
41
  Send the decide prompt with `enable_thinking: false` and a JSON-schema response format
42
  (the model is trained to answer with the decision JSON only). Full system prompt and
43
+ prompt format live in the [Agent Society repository](https://github.com/sebastianaldrin/agent-society).
44
 
45
  On CUDA, add `--flash-attn off` — llama.cpp's `auto` default silently corrupts this
46
  hybrid-SSM architecture: the JSON shape survives but the words inside turn to noise.
 
53
  - The prompt isn't cacheable on this arch, so it re-reads the whole prompt every tick.
54
  Slow on a weak GPU or CPU.
55
  - Measured: on 250 held-out teacher decisions it picks the teacher's action **78.0%**
56
+ of the time; the untuned base scores 56.8%. Method in the [evaluation doc](https://github.com/sebastianaldrin/agent-society/blob/main/docs/evaluation.md).
57
 
58
  ## License
59