Download docs/BUILD_HISTORY.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 19.2 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/BUILD_HISTORY.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/BUILD_HISTORY.md
-
curl -L -o BUILD_HISTORY.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/BUILD_HISTORY.md
Build history
How bankML was built, phase by phase, with the evidence each step stood on. This record was the crate documentation of
bankML/bankml.rs from the first commit (0.0.1, 2026-09-26) through 0.3.5 (2026-10-03). It moved here at 0.3.6 so
the crate documentation can say what bankML is, while this page keeps how it became that. The text from
The original design and phase ledger on is preserved word for word, stale
status lines included: they are the record of what was true when they were written. The full per-release detail
is in CHANGELOG.md, and every release's gate record is in testing/results/.
What bankML is now, module by module, is in modules/; the design intent the phases test is the
Thesis in TECHNICAL.md, which also stands alone
in thesis.md.
Where each phase ended up (as of 2026-10-06: 0.3.6 released, 0.3.7β0.3.9 unreleased)
| phase | what it asked | where it was done |
|---|---|---|
| P0 β wrap | bankml serve in front of llama-server: the gate, the upstream bound to the verified file, a receipt on every answer |
0.0.6 (2026-09-28); the Savante UI on it from 0.0.6β0.1.0 |
| P1 β guard + receipts | the GGUF guard, the sha256 pin, one verify gate; a receipt per answer |
guard and pin 0.0.1β0.0.2; bankml_receipt on every answer since 0.0.6; signed receipts and the THOT8 leaf are still open (0.8.0 in TODO.md) |
| P2 β own the kernels | Q1_0 and Q2_0, bit-exact against ggml, then faster | AVX2 kernels 0.0.1, threads 0.0.3, memory floor and experiments 0.0.4β0.0.5, SHA-NI 0.1.8; the GPU (Vulkan, bankML's own SPIR-V) 0.2.12β0.2.14; F16 products 0.3.4. Unreleased: the GPU limiter and per-shape calibration (0.3.7); ggml's q8_0 quantizer and vec_dot_q8_0_q8_0 for the KV cache (0.3.9). NEON and AVX-512 still open |
| P3 β own the forward | the Qwen3 block, token-identical to llama.cpp | eleven steps, 0.2.1β0.2.11 (tokenizer, template, layer 0 op by op, the whole model, the ternary model, the three attention kernels, sampling); milestone 0.3.0 (2026-09-29): Savante answered by bankML's own forward pass. The Llama graph followed in 0.3.4. The q8_0 KV cache the original ledger asked for first is unreleased (0.3.9), with llama.cpp's Hadamard rotation |
| P4 β the mindX seam | an OpenAI-compatible provider mindX can route to | serve --native 0.3.0, Ollama's API 0.3.1, the C API 0.3.2, JSON mode 0.3.3, mindX's own model 0.3.4, JSON schemas and bankml create 0.3.5, the penalties 0.3.6; mindX's default engine on its VPS since 2026-10-04. Unreleased: the rest of the sampler chain (0.3.7); the context limit, slots, the host prompt cache, logprobs (0.3.8). Open: mindX's inference discovery using it as a provider (0.9.0 in TODO.md) |
| P5 β handheld | the same crate on Android and iOS | open (see TODO.md, Handheld and distributed intelligence) |
Two lines of the original ledger no longer hold, and are kept below as written: "Nothing here runs a model yet" (true
until 0.2.7, when bankML's own forward pass first generated llama.cpp's tokens), and "out of scope until measured
need: GPU backends, β¦ training" (the GPU component arrived in 0.2.12 and mindXtrain's stages in Rust in 0.2.13, each
after a measurement asked for it). The answer() stub the status line names was removed at 0.3.6: answers come from
native::Native::complete, behind serve, the C API and bankml generate.
The phases since the ledger
The ledger below ends at 0.3.5. What followed, release by release, with the evidence each step stood on. The figures come from CHANGELOG.md; the oracles are described in oracles.md Β§5fβΒ§5h.
0.3.6 (released 2026-10-04): the penalties (O2). llama-server's repeat, frequency and presence penalties over
repeat_last_n, with the prompt in the window as llama-server fills it (sampler.rs). Evidence: oracle_penalties
and oracle_penalties_8b, 56 of 56 answers token-identical on mindx-gen39, Bonsai-1.7B and Bonsai-8B each, 12 of 12
refusals each; live 85 of 85 through /v1 and /api/chat. Record: testing/results/0.3.6.txt.
0.3.7 (unreleased): the whole sampler chain; bankML measures itself; the GPU limiter; the console.
- Typical-p, top-n-Ο, XTC, dynamic temperature and DRY, as
llama-sampler.cppwrites each, in the default order (modules/sampler.md). Evidence:oracle_samplers76 of 76 on mindx-gen39 and on Bonsai-1.7B, 16 of 16 refusals each (Bonsai-8B to be recorded);oracle_std_sort876 of 876 orders against libstdc++. - A record per completion (TTFT, prompt and generation speed, energy where RAPL is readable) at
GET /bankml/metrics(modules/metrics.md). BANKML_GPU_LIMITand per-shape calibration (modules/gpu.md). Measured on the Vega 3 with Bonsai-1.7B: with one global share the card cost 18 % (10.40 β 8.51β8.58 tok/s); per shape it is neutral (10.36 off, 10.30β10.38 on), and card memory falls from 66 to about 13 MB.- The bankML persona and console (modules/console.md), checked by
testing/test_console.py.
0.3.8 (unreleased): the serving contract. Each against llama-server b11192's answers on Bonsai-1.7B, except where
oracles.md Β§5g says the engine is its own reference: the context limit (8 of 8), slot save, restore and
erase (19 of 19), one slot with llama-server's host prompt cache (14 of 14,
modules/prompt_cache.md), and logprobs on /v1, streamed or not (14 of 14).
0.3.9 (in progress): a q8_0 KV cache and a faster grammar mask.
BANKML_CACHE_TYPE=q8_0stores K and V asq8_0blocks, 53 % of the f16 cache's bytes, with llama.cpp's Hadamard rotation around the quantized cache. Evidence:oracle_ggml_b11192_q8_0_kv_kernels(4,000 rows byte-exact, 4,000 dot products bit-exact against the shipped library) andkv_oracle_live(6 of 6 answers against llama-server with--cache-type-k/v q8_0).- The whole-vocabulary mask through a trie: median 2.94 ms against 38.8 ms (13Γ) over the oracle's 1,645 masks, every mask identical to llama.cpp's by both paths (modules/grammar.md).
- 1-bit decode speed, the third piece, is still to be measured on an idle machine.
The original design and phase ledger
As it stood in bankML/bankml.rs at 0.3.5 (commit e4d54c5).
bankml.rs β the in-house Rust player for low-bit models
Status (0.0.6, 2026-09-28): P0 serve + the Savante UI;: P1 guard + pin (one verify gate) and the P2 Q1_0 and Q2_0 (ternary) kernels are native and
proven (see the checklist). Nothing here runs a model yet β answer()'s todo!() is P0/P3. The
ternary finding: llama.cpp b11192 has no x86 Q2_0 kernel (scalar C, ~49 ns per 64 weights on the
dev box); bankml's is bit-exact with it and ~9β10Γ faster, and ggml's matmuls are ~90β98 % of its wall. Created 2026-09-26 on the operator's instruction: "create llama.cpp rust version
todo as bankml.rs", design goals optimization, succinct, verified response.
Doctrine (/core): Python prototypes, Rust ports the working architecture. The working
architecture exists as of today: PrismML Bonsai-8B (Qwen3-8B dense, ggml Q1_0, 1.16 GB,
PYTHAI/Bonsai-8B-gguf-fork) served by llama.cpp b11192 on the 2-core mindX node,
measured by scripts/ternary_diagnostics.py (data/monitoring/ternary_diagnostics.jsonl):
5/5 verified Β· 148 prompt β 130 completion tokens in 58.7 s Β· gen 2.73β3.36 tok/s (median 3.0)
prompt eval 1.1β4.4 tok/s Β· first token 0.34β8.8 s Β· peak RSS 1,786 MB Β· 2 threads, CPUQuota 150 %
One core (mindx-bonsai.service, 1 thread, CPUQuota 100 %), the same five prompts, 2026-09-26:
5/5 verified Β· 148 β 138 tokens in 105.8 s Β· gen median 2.62 tok/s Β· first token 8.9β12.3 s
97.3 CPU-seconds total = 0.705 CPU-s per completion token Β· peak RSS 1,785 MB
That row is the bar. bankml.rs is not done until it matches it on the same prompts, then beats it.
The engineer is the bankml subagent (.claude/agents/bankml.md): prefill first (the measured wall),
then memory traffic per token, with strategies from llama.cpp, vLLM, Ollama and KoboldCpp.
The three design goals, made testable
- Optimization β for a 1-bit weight the matmul is a signed sum:
y = s Β· (Ξ£ x[w=+1] β Ξ£ x[w=β1])per 128-weight group. No multiplies on the weights (ggml-exactness pins the activations to q8_0, so the sums run asmaddubs/maddby Β±1 or {0,1}). Targets on the reference node: gen β₯ 3.0 tok/s (parity), then β₯ 4.5 tok/s; prompt eval β₯ 2Γ b11192 (batched GEMM, not GEMV); RSS β€ weights + KV(ctx, kv_type) + 150 MB. Measured by ternary_diagnostics, never quoted. - Succinct β one crate, no Python, no runtime downloads. Budget: β€ 3,000 lines for
GGUF + Qwen3 forward + Q1_0/Q2_0_g64 kernels + sampler + OpenAI-compatible server. Every
dependency justified in
Cargo.tomlcomments. A single static binary. - Verified response β every answer carries its evidence, and a model that cannot be
verified does not answer:
- before load: the GGUF guard (port of minaiml
gguf_guard.py) βplay | refuse | need_more; refuse is shown with its reason, never a silent fallback (Bonsai 2 Q2_0 loads on mainline and answers in gibberish β the worst failure); - the file's sha256 must match a pinned record (
FORK.jsonof the PYTHAI fork); - kernels are proven bit-exact against ggml on recorded tensors before they are trusted;
- each response returns a
Receipt: model sha256, guard verdict, prompt/completion tokens, timings, and an optional THOT8 ternary commitment of the output (seethot8_cpu.pyβ same codec, same Keccak leaf asTHOTLib.sol), so a response can be checked later by anyone holding the model file.
- before load: the GGUF guard (port of minaiml
Phases (each ends in a measured row, or it is not done)
- P0 β wrap. (0.0.6) Not FFI:
bankml serve(serve.rs), a loopback gateway to llama.cpp b11192'sllama-server, OpenAI-compatible/v1/chat/completions;verifyin front, the upstream bound to the verified path, a receipt (with the answer's sha256) on every answer. Laptop: Savante turn 316 + 28 tokens in 125 s (prefill 2.8 tok/s β llama.cpp's). The Savante UI (sAGI/savante.py, view / interact) sits on it. - P1 β guard + receipts native. (guard and pin proven; receipts not yet emitted)
- GGUF v3 header parse + the three traps + kv_f16_bytes_per_token (
gguf.rs). Evidence: the 9 cases oftest_gguf_guard.pyas Rust tests + fail-closed extras;testing/guard_agree.py= 20/20 JSON-identical with the Python guard (realBonsai-1.7B-Q1_0.ggufsha2563d7c6c90β¦+ all synthetic cases, both engines): play Β· qwen3 Β· F32 113 / Q1_0 197 Β· 114,688 kv bytes/token. - sha256 (FIPS vectors; equals
sha256sumand the HF LFS oid on the real file) + FORK.json pin (bankml pin FILE --fork FORK.json; parses the real PYTHAI/Bonsai-8B-gguf-fork FORK.json β284a335aβ¦, refuses a mislabelled file naming both hashes). - 0.0.2:
verify= guard then pin as one gate (bankml verify FILE --fork FORK.json --json); on the real Ternary-Bonsai-8B + its FORK.json β play,e17b298dβ¦. Guard hardened against hostile headers (nested arrays refuse instead of overflowing the stack; KV-size overflow refuses). - Receipt emitted per answer, THOT8 leaf (needs P0/P3 to have an answer to sign).
- GGUF v3 header parse + the three traps + kv_f16_bytes_per_token (
- P2 β own the kernels. (Q1_0 and Q2_0 proven on x86 AVX2; NEON, real activations open)
- Q1_0 exactly as ggml b11192 defines it (
q1_0.rs, read from source): f16dfirst, 16 sign bytes LSB-first; the matmul isggml_vec_dot_q1_0_q8_0against q8_0 activations (AVX2 quantizer: 127/amax, round-half-even). Evidence (oracle_ggml_b11192_real_bonsai_1_7b, against the exported symbols of the sha256-checked b11192 ubuntu-x64 release,libggml-cpu-haswell.so): all 197 Q1_0 tensors = 1,719,904,256 weights dequantized bit-exact; 788/788 q8_0 rows byte-exact; 788/788 vec_dot results bit-exact vs ggml AVX2, and 788/788 vs ggml generic (whose shipped build fuses to FMA β read from the disassembly). AVX2 == scalar model on 3,500 + 2,000 randomized cases. - Measured (dev box Ryzen 3 3200U, Zen+, 1 thread, in-process A/B vs ggml's own kernel, same bits):
decode GEMV 12288Γ4096 parity (min 10.0 vs 10.0β10.2 ns/block;
vec_dot_act, activation scales converted once per token); prefill 1Γ4 tile 1.10Γ ggml's per-pair vec_dot (min 9.1 vs 10.1 ns/(blockΒ·col)). Not the node: a Zen3 row is the deciding measurement. - Q2_0 (id 42, group 64) exactly as ggml b11192 defines it (
q2_0.rs, read from the tag's source): f16dfirst, 16 bytes of 2-bit codes LSB-first, code c β (c β 1)Β·d β {β1, 0, +1, +2} (the Bonsai file never uses +2: 31 % / 38 % / 31 % / 0 %; the kernels take it anyway).vec_dot_typeq8_0. On x86, b11192 has no Q2_0 kernel:arch-fallback.hrenames the generic C toggml_vec_dot_q2_0_q8_0, no repack, no sgemm case; the haswell symbol is scalar (64imul/block, read from the disassembly), used per (row, column) for decode and prefill. Evidence (oracle_ggml_b11192_real_ternary_bonsai_8b, real Ternary-Bonsai-8B sha256e17b298dβ¦, via mmap): all 254 Q2_0 tensors = 8,188,239,872 weights dequantized bit-exact; 762/762 q8_0 rows byte-exact; 762/762 vec_dot bit-exact vs ggml haswell (scalar model, AVX2, portable path), and 762/762 of a no-FMA model vs the baselinelibggml-cpu-x64.so(which differs from haswell in 19/762: the float order is really being tested). 3,300 randomized cases incl. code 3, q = β128, tail blocks. - Measured (dev box, 1 thread, in-process A/B vs ggml's own Q2_0 kernel, real layer-0 weights,
alternating, 4 runs; the laptop's boost clock moves both, so ratios are the result): decode GEMV
9.1β10.6Γ (ggml 48.6β55 vs bankml 4.9β6.0 ns/block min;
output151669Γ4096: 492 β 51 ms); prefill 1Γ4 tile 12.3β13.1Γ (49 vs 3.96 ns/(blockΒ·col)). The per-token activation layout costs 11 Β΅s (n = 4096) / 29 Β΅s (n = 12288), β 2 ms per token: outside the matmul timings, 0.5 % of them. The kernel: two blocks per 256-bit register, codes by(v >> 2r) & 3,maddubs(c, q) β Ξ£q, the activation (not the weights) permuted once per token, ggml's float chain kept serial. One-token budget (decode_budget_q2_0, all 253 matmuls, 2.13 GB, each tensor once in file order, llama-bench run before and after at the same clock): 3 threads β llama-bench tg 2.59β2.68 s/token, ggml's matmuls 2.54β2.64 s (β 98 % of the wall), bankml 0.36β0.37 s (7.0Γ; matmul-only ceiling 2.77 tok/s); 1 thread β llama-bench 4.50β5.21 s, ggml matmuls 4.22β4.42 s, bankml 0.49 s (8.7Γ; ceiling 2.06 tok/s). Not the node: a Zen3 row decides. - 0.0.3 threads:
par::Pool(persistent, zero-dependency) +mat_vec_par/mat_mul_par, bits independent of thread count. One ternary token's 253 matmuls: 0.23β0.25 s at 3 threads (ggml 2.27β2.36 s), below ggml's 1-bit 0.34 s; 1-bit at parity. New 8B Q1_0 oracle: 254 tensors bit-exact. - 0.0.4: memory floor 15β17 GB/s (both kernels compute-bound: ternary ~2Γ, 1-bit ~5.5Γ above
it);
mat_mul_act1-bit prefill 1.27β1.33Γ ggml; two bit-exact decode variants measured and rejected. - 0.0.5: three ternary experiments bit-exact and not reliably faster (repacked quads 0.90β1.075Γ, two-row tile 0.80Γ, prefetch Β±3 %); no kernel change. Both kernels are at the Zen+ instruction limit: the next gains are P3 and a Zen3 row.
- NEON; AVX-512 (the node has none); bit-exact on dumped real activations (today: real weights Γ synthetic activations); a Q1_0 GEMM that beats 1.1Γ for prompt eval (the Q2_0 layout above runs at 5.0 ns per 64 weights vs Q1_0's 7.2 β porting it to Q1_0 is the next experiment).
- Q1_0 exactly as ggml b11192 defines it (
- P3 β own the forward. Qwen3 dense block (RMSNorm, RoPE, GQA, SwiGLU) with a quantized KV cache (q8_0 first β at 1.125 bpw the KV cache, not the weights, sets the RAM ceiling). Drop the FFI. Parity with P0 on all 5 prompts at temperature 0 (token-identical is the goal).
- P4 β mindX seam. Serve on loopback; register as an OpenAI-compatible provider in mindX
(the vLLM handler shape);
inference_budget.record_inference(..., model=, caller=)sees real usage;/substrate/and feedback show its row. Then the 8B migration can route to it.- 0.3.1: Ollama's API on
serve --native(ollama.rs), a registry of pinned models, one resident, verified on every load,keep_alive; native only (docs/OLLAMA.md: what mindX asks of Ollama, and the O1βO8 track). - 0.3.2: a C API (
capi/,libbankml.so/.a,capi/include/bankml.h, docs/CAPI.md) β open (the same verify), chat (the same answer and receipt asserve --native), and a printf-stylebankml_logdefined in Rust as a C-variadic function (Rust 1.99); the llama.h-shaped seam for embedding bankML in another program. - 0.3.3: JSON mode (O6's first cut,
grammar.rs) β llama.cpp's GBNF engine ported, the grammar llama-server builds forresponse_format: json_objectwith its prefill, and the sample β check β masked redraw ofcommon_sampler_sample;/v1response_format/grammar, Ollamaformat: "json",bankml_chat,bankml generate --json, token-identical to llama-server b11192 greedy and seeded. - 0.3.4: O4 β tied embeddings (Bonsai-1.7B native), F16 weights (
f16.rs: ggml'svec_dot_f16and llamafile's tinyBLAS, chosen by shape as ggml chooses them), the Llama graph (NORM RoPE, no Q/K norms), SmolLM2'ssmollmpre-tokenizer and two ChatML templates: SmolLM2-135M-Instruct and mindX's ownmindx-gen39served natively, token-identical to llama-server b11192 (greedy, seeded,/v1,/api, C API, JSON mode). - 0.3.5: O6b β JSON schemas (
schema.rs: llama.cpp'sjson_schema_to_grammarand the chat parser's wrapping per template), answers token-identical to llama-server b11192 on all five native models, the content rule against llama.cpp's own parser; O5's first cut βbankml convert(byte-identical to b11192's converter) andbankml create(derived models as verified layers), mindX's persona layer token-identical end to end.
- 0.3.1: Ollama's API on
- P5 β handheld. Same crate β Android (NDK) / iOS as the minaiml BROBOT engine alternative.
Out of scope until measured need: GPU backends, PrismML fork types (PQ2_0, PTQ1_0), Bonsai 2's Hadamard basis (watch ggml-org/llama.cpp#27779), training.