azharmo's picture
Upload RESULTS.md with huggingface_hub
f1bfbc3 verified
|
Raw
History Blame Contribute Delete
2.59 kB

RESULTS — real run on a laptop CPU (no GPU)

Machine: 8-core x86_64, 15 GB RAM, PyTorch 2.14 CPU, Python 3.12. Model: ~3.1M params, d_model=128, 3 layers, 4 heads, word-level vocab of ~20,749. Data (subsampled for a toy): AG News 1500 + BoolQ 1500 + SST-2 1500 → 5700 examples. Training: 3 epochs, AdamW lr=3e-4, proper-scoring-rule losses.

Training loss

epoch loss
1 1.5539
2 1.2664
3 1.0011

Evaluation (held-out, built from public test/validation splits)

noul head (BoolQ yes/no + SST-2 sentiment) — 2000 eval rows

  • accuracy = 0.597
  • Brier score = 0.236
  • Expected Calibration Error (ECE) = 0.0265 (excellent for a toy: predicted confidence tracks actual accuracy within ~2.7%)

choice head (AG News 4-way routing) — 1000 eval rows

  • accuracy = 0.759
  • Brier = 0.331
  • top-1 basically always the argmax

temperature scaling

  • best noul temperature T = 0.8 → ECE 0.0254 (slight gain over raw)

Honest interpretation

  • The accuracy is low (59.7% noul, 75.9% choice) — expected: ~3M params, ~9% of AG News train, 25% of BoolQ, 2% of SST-2, only 3 epochs, no GPU. This is an architecture demo, not a SOTA model.
  • The calibration is genuinely good (ECE ~2.7%), which is the architecturally interesting result: a proper-scoring-rule loss + temperature scaling makes the model honest about its own confidence, which is the whole point of a "System One" decision model.
  • The demo (one state, three parallel typed questions) runs and returns Jev-shaped output.

Inferencing demo (screenshot: SCREENSHOT_serve.png)

One state + three typed questions, answered in parallel:

is_urgent:    noul   -> 0.6742  (true)
department:   choice -> {billing:0.0, technical:1.0, sales:0.0}, argmax=technical, conf=1.0
frustration:  score  -> 0.518

Plus a news-routing choice demo (Sci/Tech 0.74).

Parallel-encoding benchmark (proves the state is encoded ONCE)

Full answer() time on ONE state as the question count grows (CPU, single request):

questions ms
1 27.3
3 35.7
6 46.3
12 101.0
1→12 questions costs ~3.7x, not ~12x: the state encoder runs once; growth is only the
question-batch pass. This is the measurable parallel property (not a fake "parallel" from
a fine-tuned generator). See ARCHITECTURE.png.

How to reproduce

source .venv/bin/activate
python -m jev_toy.train --epochs 3 --agnews 1500 --boolq 1500 --sst2 1500
python -m jev_toy.eval --ckpt checkpoints/model.pt
python -m jev_toy.serve --ckpt checkpoints/model.pt