DeskMind Brain 0.8b · 得心

得心,应手。 Brain is the decision model of DeskMind, open-source models and tools that let an agent see your screen, decide the next step and act on your own Mac. It answers System One–style typed questions (which operation next, which element, is the goal met, should I ask the user first) and returns calibrated probabilities read from answer-letter logits: nothing is generated or parsed. The server is wire-compatible with POST /v1/systemone.

Try it and read more: the DeskMind Mac app runs this model on your Mac · how the decision models were trained · code, app and bench

This model is the fast tier of the Brain router: it answers every step first, and hands the steps it is unsure of, or that are costly to get wrong, to brain-4b.

  • Release: G18b, revision g18b-q8 (8-bit MLX, prompt format 3). main follows the current release.
  • Routing threshold: shipped in this model's deskmind.json (router_threshold: 0.96); deskmind-brain-serve reads it, and --threshold overrides it.
  • Download: 0.8 GB.
  • Code and docs: deskmind-ai/brain. Keep deskmind.json next to the weights: it records the prompt format the model was trained with.

Revisions

Revision What it is
g18b-q8 Current release. G18b, router threshold 0.96; used by the DeskMind app v0.3.0.
g17-q8 Test build, not a release.
g14-q8 Earlier release (G14, prompt format 3, router threshold 0.94); used by the DeskMind app v0.2.0.
g13-q8 Earlier release (G13, prompt format 3).
v7b-q8 First published build (checkpoint G11b, prompt format 2).

Use

git clone https://github.com/deskmind-ai/brain && cd brain
uv sync --extra mlx
uv run hf download deskmind/brain-0.8b --revision g18b-q8 --local-dir models/brain-0.8b
uv run hf download deskmind/brain-4b --revision g18b-q8 --local-dir models/brain-4b
uv run deskmind-brain-serve --predictor mlx:models/brain-0.8b --escalate-to mlx:models/brain-4b \
  --two-stage --port 8796

The 4B is 4.5 GB to download and the 0.8B 0.8 GB. If hf download fails with CAS Client Error, retry with HF_HUB_DISABLE_XET=1 in front of the command. In mainland China, ModelScope carries the same files:

uvx modelscope download --model gxcsoccer/brain-0.8b --revision g18b-q8 --local-dir models/brain-0.8b

Each reply carries a routing record: who answered, and why. Full walkthrough: deskmind-ai/brain.

Results

All numbers are our own runs; method and full tables are in docs/results.md.

Real macOS desktop, bench suite v25, 13 sandbox tasks × 3 runs, strict pass, run through the DeskMind app on an M4 Pro (48 GB):

config pass false "done" decision time p50 / p95
Router G18b (0.8B → 4B, 8-bit, threshold 0.96) 39/39 0 2.85 / 9.82 s (208 decisions)
Router G14 (earlier release, threshold 0.94) 36/39 0 0.57 / 5.25 s
  • Steps the 0.8B answers itself take p50 0.48 s / p95 0.66 s; steps escalated to the 4B take p50 3.6 s / p95 9.8 s. About 70% of steps escalate.
  • The G18b runs had the app's optional checks and notes off; there were no environment errors and no no-progress loops.
  • 39 runs over 13 tasks is a small sample, and runs cluster by task (a task tends to pass 3/3 or 0/3).
  • On the earlier suite v23, the G14 router passed 35/38 (92%) and Jev (TypeSafe, hosted) 33/38 (87%). There is no Jev run on v25.

JevBench v1.4.2, public set (231 items, the board's public_accuracy column), run locally with the official jevbench.cli:

easy (48) original (72) hard (111) public (231)
Brain 4B, G18b 48 69 76 0.835
Brain 0.8B, G18b 48 56 63 0.723
Router G18b (threshold 0.96) 48 65 71 0.797
Brain 4B, G14 48 67 85 0.866
  • Public items only. The board's headline JevBench Score also weighs 308 sealed items, calibration, speed and cost; we have not been scored on the sealed set.
  • Calibration of the G18b 4B: Brier 0.269, ECE 0.089 (G14: 0.230, 0.079).
  • Contamination check: the G18b training mix (100,703 items) shares no word 13-gram and no option set with the 231 public items (0/231).

Limitations

  • Speed: the G18b 0.8B's confidences sit in a narrow band (about 0.94–0.97), so at threshold 0.96 most steps go to the 4B and a typical decision takes about 3 s, slower than hosted models.
  • General judgement: against G14, the 4B dropped on JevBench's hard tier (85 → 76 of 111), mostly temporal/numeric, hard-judgement and multi-hop items. G18b was kept for its real-desktop reliability.
  • Saying "done": on the Chinese exact-text task the file was right in all 3 runs, but the model never said "done" and used the full 20-step budget. The grader checks the final state, so these count as passes.
  • Scope: trained and tested on macOS Finder and TextEdit sandbox tasks plus web and form decisions; untested elsewhere.
  • Probabilities are not guarantees: a probability is the model's own weighting of the options, not proof that the step is right.

Training

LoRA distillation on Qwen/Qwen3.5-0.8B (KL to teacher distributions plus cross-entropy to labels), merged and quantized to 8 bits. Desktop data comes from DAgger in sandboxed macOS tasks, with every visited state labelled by a desktop oracle, plus counterexamples that break label shortcuts. Web, form and evidence items come from public datasets and synthetic tasks, generated and labelled with hosted frontier-model teachers. No Jev outputs were used as labels, and the data contains no real user data. Details: docs/training.md.


中文: 得心(DeskMind)Brain 的快速档:每一步先由它回答,没把握或容易出大错的步骤交给 4B。路由门槛 0.96 写在本模型的 deskmind.json 里。给电脑操作 agent 的每一步做带类型、带把握程度的决策,用 MLX 在 Apple Silicon 本地运行。

  • 当前发布版: G18b,版本 g18b-q8(8 位 MLX);DeskMind app v0.3.0 使用此版本。
  • 真机成绩: bench v25,13 个沙箱任务各跑 3 轮,通过 DeskMind app 运行,39/39 通过,没做完就说完成 0 次。0.8B 直接回答的步骤中位 0.48 秒;约 70% 的步骤交给 4B,中位 3.6 秒。
  • JevBench v1.4.2 公开题(231 道): 4B 0.835,0.8B 0.723,路由 0.797;训练数据与公开题无重合(0/231)。
  • 下载: hf download 加 --revision g18b-q8;国内可用 ModelScope: uvx modelscope download --model gxcsoccer/brain-0.8b --revision g18b-q8 --local-dir models/brain-0.8b。
  • 详见 deskmind-ai/brain。

License

Apache-2.0 (see LICENSE and NOTICE). Fine-tuned from Qwen/Qwen3.5-0.8B (Copyright Alibaba Cloud, Apache-2.0). The DeskMind name, 得心 and the logo are not covered by this licence.

Downloads last month
51
Safetensors
Model size
0.8B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for deskmind/brain-0.8b

Quantized
(305)
this model

Collection including deskmind/brain-0.8b