SavantebankML
Δ deltaverse · cryptoAGI

bankML

Verified 1-bit and ternary language models on the CPU you already have: exactly the same answers as llama.cpp, checked bit for bit, then faster.

Updated 7 October 2026: bankML 0.4.0, the milestone, is the latest release.

Latest release: v0.4.0 — native serving complete

Released 7 October 2026, after one run of the full release gate in which every stage passed. Everything Savante and mindX ask of llama-server is now answered by bankML's own engine, identical to llama-server b11192.

the release · what changed, in full · the gate record · the speed test, with the machine's load

The thesis — Professor Codephreak and Gregory L. Magnusson

bankML is built to test its authors' design intent, stated in their own words over the course of the project. Every release is measured against it. Exact and precise first, then efficient and optimized: the same bits as the reference, every value held to its last bit and every measurement stated with its resolution; then the CPU you already have, used well, and speed that counts only once the oracle passes.

  1. Port what works, and only what works. "We prototyped in python and we switch to rust from working architecture." A proven system's behaviour is the specification, so exactness against its compiled code comes first.
  2. Optimization, succinct, and verified response. Speed is measured, never quoted. The runtime is one crate with no dependencies. A model that cannot be verified does not answer.
  3. The standing machine is the CPU you already have. One server powers one autonomous AI system for about a dollar a day. Accelerators are rented for events, not owned.
  4. Consume the inference, and prefer the better answer. "A better answer is more important than speed where speed under one minute is fine for now". That is why the better ternary model is the one made fast.
  5. Know the limit, and code around it. The guard refuses, with its reason, a file that would answer in nonsense. Where the reference is slow, the answer is a new kernel, not acceptance.
  6. Improve by increments, with the oracle in the loop. A speed counts only when every oracle has passed on the same code; the rejected variants are kept as evidence.
  7. Use the machine at hand. Every number was taken on a two-core laptop. Where the machine disturbs a measurement, the record says so.
  8. Prove the data; keep the data. What travels is a commitment (a sha256, a CID, a Merkle root), never the data itself.

the Thesis, in full, with its sixteen contributions · the argument, with the contemporary field · hear it read

Ask bankML

Start your own bankML for this page (Linux, AVX2; free)
  1. git clone https://github.com/cryptoAGI/bankml && cd bankml && ./install.sh --space — builds bankML, imports and verifies Bonsai-8B, and starts it so this page () may talk to it. Already installed? git pull && ./install.sh build start --space
  2. Press connect. Your browser may ask to let this page reach devices on your local network: that is the request for your own bankML on 127.0.0.1, and nothing else.
  3. ./install.sh start --no-space takes the permission back.

bankML in one paragraph

bankML runs language models on the computer you already have: a laptop, a small server, no graphics card needed. It specialises in 1-bit and ternary models, which store each weight in one or two bits instead of sixteen. An 8-billion-parameter model then fits in 1.2 GB (1-bit) or 2.3 GB (ternary). bankML gives exactly the same answers as llama.cpp, the reference engine, and checks that it does, bit for bit, before it is allowed to be faster. Every answer carries a receipt that says which model file produced it.

9.4–10×faster per ternary matrix than llama.cpp
≈ 8×whole ternary answers vs llama-server
0external dependencies
bit-exactagainst llama.cpp b11192

the full story: README · the technical report: TECHNICAL.md

bankML, in its own words

The console that talks to bankML speaks as bankML itself, from its persona file, sAGI/personas/bankml.persona. It knows its own use only from measurement: each question carries a SELF block of what bankML measured at that moment, and a value it did not measure it says it did not measure.

“The same bits first, then the speed.”

Its oath. I answer only from verified weights. What I cannot reproduce exactly I refuse, with the reason. Every answer carries its receipt, and what I say about my own use I have measured.

the persona's sources: README's design goals and the thesis · the console (sAGI/console.py)

This Space is the source

Every file of bankML is here: the Rust engine (bankML/), the C API (capi/), the oracles and the release gate (testing/), the tools, the docs and Savante's UI (sAGI/) — browse the files. The source of record is github.com/cryptoAGI/bankml.

Talk to bankML yourself. On a Linux machine with AVX2:

git clone https://github.com/cryptoAGI/bankml && cd bankml && ./install.sh
then python3 sAGI/console.py and open http://127.0.0.1:7875.

the live engine here: this Space already carries its Dockerfile and hf/start.sh (bankML serving the ternary Bonsai-8B, the console in front, public and read-only); it runs once the Space is switched to Docker hardware

The advantages

1. It is fast where it matters most

llama.cpp has no fast x86 code for ternary weights; it falls back to plain C. bankML wrote the missing kernel. Each matrix is 9.4–10× faster, and whole answers on the ternary model come out at about 8× llama-server's speed (2.3–2.4 tokens/s against 0.30 on the same laptop). The ternary model gives the better answers of the two, so this is the speed that counts.

PERFORMANCE.md · the kernel: modules/q2_0.md, bankML/q2_0.rs

2. It is exact, and proves it

Speed only counts if the numbers are the same. Every kernel, the tokenizer, the chat templates, the samplers and whole conversations are checked against llama.cpp b11192's own compiled code. These checks are called oracles, and a release ships only when every oracle passes in its release gate.

oracles.md · the gate: testing/release_gate.sh and its records in testing/results/

3. It refuses to answer from a model it cannot verify

Before a model answers, three gates run. A guard reads the file's header and refuses known-bad files. A pin checks the file's sha256 against the record of where it came from. The oracle, run ahead of time, has proven the arithmetic. Each answer then carries a receipt with the model's sha256 and the hashes of the request and the answer.

TECHNICAL.md §III.1 · receipts (usage.md §9) · the guard: bankML/gguf.rs · the pin: bankML/sha256.rs

4. It is small and has nothing to install around it

bankML is one Rust crate with zero external dependencies. The GGUF reader, sha256, f16 maths, the thread pool, the HTTP server and every kernel are written in the crate. You get one binary and nothing else to trust.

one page per source file · Cargo.toml

5. It speaks the languages your tools already use

6. It measures itself

bankML reports its own time to first token, tokens per second and, where the machine allows, the energy each token costs. It also limits how much of a graphics card it may use.

metrics.md · console.md · the GPU: gpu.md

What people use it for

A private assistant on your own machineSavante, a chat page on 127.0.0.1:7873, answered by bankML — usage.md §4, playback.md
The engine behind mindXmindX's default model server on its VPS since 4 October 2026, in place of Ollama — OLLAMA.md
A drop-in Ollama or OpenAI serverbankml serve --native --registry — usage.md §6a, install.md
Your own model, verifiedbankml convert (safetensors → GGUF, byte-identical to llama.cpp's) and bankml create (a persona layer from a Modelfile) — convert.md, create.md
Structured answers for programsJSON mode, JSON schemas and grammars — usage.md
Inference inside another programlibbankml and bankml.h — CAPI.md
Agents with a provable historyreceipts, Merkle commitments of the conversation, THOT bundles and iNFTs — usage.md §8a, §8e
Meaning search over notesbge-m3 embeddings fused with keyword search — embedding.md

new here? start with ./install.sh: usage.md §1

How Rust helps bankML go fast and stay exact

Rust lets bankML write code as close to the hardware as C, while the compiler checks much of what C leaves to the programmer.

Where it is going

The road is laid out release by release in TODO.md, and each step counts only once its oracle passes.

Our hope for bankML

We hope bankML shows that a capable, private AI does not need a data centre or a graphics card: that an eight-billion-parameter model can run on the laptop in front of you and give answers you can check. We want every speed claim to come with proof that the numbers are the same, so that "faster" always means faster and correct. We hope the ternary kernel goes back upstream to llama.cpp, so everyone benefits from it. We want bankML to become the dependable, verifiable engine under Savante, mindX and the agents built on them. And we want it to stay small enough that one person can read all of it.

the authors' own words: the Thesis · the argument in full, with the contemporary field: the thesis · among other engines and papers: research.md · how it was built: BUILD_HISTORY.md