bankML
Verified 1-bit and ternary language models on the CPU you already have: exactly the same answers as llama.cpp, checked bit for bit, then faster.
Updated 7 October 2026: bankML 0.4.0, the milestone, is the latest release.
Latest release: v0.4.0 — native serving complete
Released 7 October 2026, after one run of the full release gate in which every stage passed. Everything Savante and mindX ask of llama-server is now answered by bankML's own engine, identical to llama-server b11192.
- 1-bit at llama-server's speed. In the pinned 8B test, bankML was at least as fast in all three rounds (median 2.02 against 0.77 tokens per second under the same load), and every answer was identical. The default engine is now bankML's own for 1-bit as well as ternary models.
- A smaller conversation memory. A
q8_0cache with llama.cpp's own Hadamard rotation: 53 % of the memory, the same answers (6 of 6). - The serving contract. The context limit, saved and restored slots, a prompt cache for conversations that take turns, and token probabilities, each identical to llama-server's.
- It shows its work. A status page on the engine's own address (CPU, memory, disk, GPU, its
log). The console opens on the ultimate input field and has Engine, Diagnostics and Thesis tabs. A
scientific diagnostic writes every token's probability from both engines to 18 decimals: 32 of 32
bit-equal, difference
0.000000000000000000. - Kernels on this laptop, three threads: 1-bit level with llama.cpp's own (1.02×), ternary 8.9× faster, the same bits.
the release · what changed, in full · the gate record · the speed test, with the machine's load
The thesis — Professor Codephreak and Gregory L. Magnusson
bankML is built to test its authors' design intent, stated in their own words over the course of the project. Every release is measured against it. Exact and precise first, then efficient and optimized: the same bits as the reference, every value held to its last bit and every measurement stated with its resolution; then the CPU you already have, used well, and speed that counts only once the oracle passes.
- Port what works, and only what works. "We prototyped in python and we switch to rust from working architecture." A proven system's behaviour is the specification, so exactness against its compiled code comes first.
- Optimization, succinct, and verified response. Speed is measured, never quoted. The runtime is one crate with no dependencies. A model that cannot be verified does not answer.
- The standing machine is the CPU you already have. One server powers one autonomous AI system for about a dollar a day. Accelerators are rented for events, not owned.
- Consume the inference, and prefer the better answer. "A better answer is more important than speed where speed under one minute is fine for now". That is why the better ternary model is the one made fast.
- Know the limit, and code around it. The guard refuses, with its reason, a file that would answer in nonsense. Where the reference is slow, the answer is a new kernel, not acceptance.
- Improve by increments, with the oracle in the loop. A speed counts only when every oracle has passed on the same code; the rejected variants are kept as evidence.
- Use the machine at hand. Every number was taken on a two-core laptop. Where the machine disturbs a measurement, the record says so.
- Prove the data; keep the data. What travels is a commitment (a sha256, a CID, a Merkle root), never the data itself.
the Thesis, in full, with its sixteen contributions · the argument, with the contemporary field · hear it read
Ask bankML
Start your own bankML for this page (Linux, AVX2; free)
git clone https://github.com/cryptoAGI/bankml && cd bankml && ./install.sh --space— builds bankML, imports and verifies Bonsai-8B, and starts it so this page () may talk to it. Already installed?git pull && ./install.sh build start --space- Press connect. Your browser may ask to let this page reach devices on your local network: that is the request for your own bankML on 127.0.0.1, and nothing else.
./install.sh start --no-spacetakes the permission back.
bankML in one paragraph
bankML runs language models on the computer you already have: a laptop, a small server, no graphics card needed. It specialises in 1-bit and ternary models, which store each weight in one or two bits instead of sixteen. An 8-billion-parameter model then fits in 1.2 GB (1-bit) or 2.3 GB (ternary). bankML gives exactly the same answers as llama.cpp, the reference engine, and checks that it does, bit for bit, before it is allowed to be faster. Every answer carries a receipt that says which model file produced it.
the full story: README · the technical report: TECHNICAL.md
bankML, in its own words
The console that talks to bankML speaks as bankML itself, from its persona file,
sAGI/personas/bankml.persona.
It knows its own use only from measurement: each question carries a SELF block of what bankML measured at that moment, and a value it did not measure it says it did not measure.
“The same bits first, then the speed.”
Its oath. I answer only from verified weights. What I cannot reproduce exactly I refuse, with the reason. Every answer carries its receipt, and what I say about my own use I have measured.
the persona's sources: README's design goals and the thesis ·
the console (sAGI/console.py)
This Space is the source
Every file of bankML is here: the Rust engine (bankML/), the C API (capi/), the oracles and
the release gate (testing/), the tools, the docs and Savante's UI (sAGI/) —
browse the files. The source of record is
github.com/cryptoAGI/bankml.
Talk to bankML yourself. On a Linux machine with AVX2:
git clone https://github.com/cryptoAGI/bankml && cd bankml && ./install.sh
then python3 sAGI/console.py and open http://127.0.0.1:7875.
the live engine here: this Space already carries its Dockerfile and
hf/start.sh (bankML serving the ternary Bonsai-8B, the console in front, public and read-only);
it runs once the Space is switched to Docker hardware
The advantages
1. It is fast where it matters most
llama.cpp has no fast x86 code for ternary weights; it falls back to plain C. bankML wrote the missing kernel. Each matrix is 9.4–10× faster, and whole answers on the ternary model come out at about 8× llama-server's speed (2.3–2.4 tokens/s against 0.30 on the same laptop). The ternary model gives the better answers of the two, so this is the speed that counts.
PERFORMANCE.md ·
the kernel: modules/q2_0.md,
bankML/q2_0.rs
2. It is exact, and proves it
Speed only counts if the numbers are the same. Every kernel, the tokenizer, the chat templates, the samplers and whole conversations are checked against llama.cpp b11192's own compiled code. These checks are called oracles, and a release ships only when every oracle passes in its release gate.
oracles.md · the gate:
testing/release_gate.sh
and its records in testing/results/
3. It refuses to answer from a model it cannot verify
Before a model answers, three gates run. A guard reads the file's header and refuses known-bad files. A pin checks the file's sha256 against the record of where it came from. The oracle, run ahead of time, has proven the arithmetic. Each answer then carries a receipt with the model's sha256 and the hashes of the request and the answer.
TECHNICAL.md §III.1 ·
receipts (usage.md §9) ·
the guard: bankML/gguf.rs ·
the pin: bankML/sha256.rs
4. It is small and has nothing to install around it
bankML is one Rust crate with zero external dependencies. The GGUF reader, sha256, f16 maths, the thread pool, the HTTP server and every kernel are written in the crate. You get one binary and nothing else to trust.
one page per source file ·
Cargo.toml
5. It speaks the languages your tools already use
- the OpenAI chat API (
/v1/chat/completions), streamed or not — serve.md - the Ollama API (
/api/chat,/api/generate, a registry of pinned models), so programs written for Ollama work unchanged — OLLAMA.md - a C library (
libbankml) for C, C++, Python, Go or Swift programs — CAPI.md - JSON mode and JSON schemas: answers forced into a shape you choose — grammar.md, schema.md
6. It measures itself
bankML reports its own time to first token, tokens per second and, where the machine allows, the energy each token costs. It also limits how much of a graphics card it may use.
metrics.md · console.md · the GPU: gpu.md
What people use it for
| A private assistant on your own machine | Savante, a chat page on 127.0.0.1:7873, answered by bankML —
usage.md §4,
playback.md |
| The engine behind mindX | mindX's default model server on its VPS since 4 October 2026, in place of Ollama — OLLAMA.md |
| A drop-in Ollama or OpenAI server | bankml serve --native --registry —
usage.md §6a,
install.md |
| Your own model, verified | bankml convert (safetensors → GGUF, byte-identical to llama.cpp's) and
bankml create (a persona layer from a Modelfile) —
convert.md,
create.md |
| Structured answers for programs | JSON mode, JSON schemas and grammars — usage.md |
| Inference inside another program | libbankml and bankml.h —
CAPI.md |
| Agents with a provable history | receipts, Merkle commitments of the conversation, THOT bundles and iNFTs — usage.md §8a, §8e |
| Meaning search over notes | bge-m3 embeddings fused with keyword search — embedding.md |
new here? start with ./install.sh:
usage.md §1
How Rust helps bankML go fast and stay exact
Rust lets bankML write code as close to the hardware as C, while the compiler checks much of what C leaves to the programmer.
- Direct access to the CPU's vector instructions. The kernels use AVX2 through
std::arch, compiled with#[target_feature(enable = "avx2,fma,f16c")]. bankML checks at run time which instructions the CPU has, so one binary runs everywhere and uses the fast path where it can. — q1_0.md, q2_0.md - No hidden costs. No garbage collector and no runtime pauses. Iterators and fixed-size
chunks (
as_chunks) compile down to plain loops, which lets the compiler drop bounds checks inside the hot kernels. - Safe threads, the same bits on any thread count. The thread pool hands out rows from an
atomic counter. Ownership rules guarantee each output chunk has exactly one writer, so threads can never
race, and the answer does not depend on how many threads ran. —
par.md,
bankML/par.rs unsafekept in small, named places. The memory map, the vector intrinsics and the GPU loader carry documented safety contracts behind safe functions; the compiler checks the rest. — gguf.md- Zero-copy model loading. The model file is memory-mapped, and the kernels read the
weights straight from the mapping. —
bankML/gguf.rs - Hostile files handled safely. All header arithmetic is checked (
checked_add,checked_mul), so a malformed file is refused with a reason instead of crashing. - Whole-program optimisation. Release builds use link-time optimisation, one
code-generation unit and abort-on-panic: one small, fast binary. —
Cargo.toml - A pinned toolchain. Rust 1.99 is pinned, so a build today gives the same machine code as
a build next month, which an exactness project needs. —
rust-toolchain.toml
Where it is going
The road is laid out release by release in TODO.md, and each step counts only once its oracle passes.
- 0.3.7 (released): llama-server's whole sampler chain, bankML measuring itself, a GPU limiter, and the bankML console.
- 0.3.8 (released): llama-server's behaviour at the context limit, saved and restored conversations (slots), a prompt cache so conversations taking turns each keep their context, and token probabilities, streamed or not.
- 0.4.0, the milestone (released 7 October 2026): native serving complete. bankML answers everything Savante and mindX ask of llama-server, and its own engine is chosen by default for both 1-bit and ternary models.
- 0.4.x: stop the work of a request whose client has gone, and a one-line measured self for the console, so long conversations keep their cache.
- 0.5.0, hardware: ARM phones and tablets (NEON), newer x86 instructions (AVX-512), and more graphics cards.
- 1.0.0: llama.cpp is needed only as the oracle in the gate, never at run time. bankML runs at llama.cpp's speed or better on every supported format, its interfaces are stable, and receipts are signed.
Our hope for bankML
We hope bankML shows that a capable, private AI does not need a data centre or a graphics card: that an eight-billion-parameter model can run on the laptop in front of you and give answers you can check. We want every speed claim to come with proof that the numbers are the same, so that "faster" always means faster and correct. We hope the ternary kernel goes back upstream to llama.cpp, so everyone benefits from it. We want bankML to become the dependable, verifiable engine under Savante, mindX and the agents built on them. And we want it to stay small enough that one person can read all of it.
the authors' own words: the Thesis · the argument in full, with the contemporary field: the thesis · among other engines and papers: research.md · how it was built: BUILD_HISTORY.md