Download docs/research.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 16.8 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/0ab4e0b6dd5a713718f7d8e1409d62164949e896/docs/research.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml@0ab4e0b6dd5a713718f7d8e1409d62164949e896/docs/research.md
-
curl -L -o research.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/0ab4e0b6dd5a713718f7d8e1409d62164949e896/docs/research.md
Research: where bankml stands
A survey of Rust-native LLM inference, 1-bit and ternary CPU inference, and verifiable inference, made on
2026-09-28/29 to answer one question: is bankml at the cutting edge, and where is it not? Every link below was
checked to resolve on those dates unless marked. bankml's own numbers are its own measurements
(TECHNICAL.md, PERFORMANCE.md, testing/results/); no third
party has benchmarked bankml.
Updated 2026-10-06 for bankml's own status only: v0.3.6 is the latest release, and 0.3.7β0.3.9 are unreleased (CHANGELOG.md). The survey of other projects is as found on 2026-09-28/29 and has not been redone; their state may have moved since. Where bankml's status changed, the original statement is kept and the change is marked.
The answer in brief
Where bankml is at or near the edge
Bit-exact reproduction of the compiled llama.cpp library, tested by an oracle, from Rust. bankml's
Q1_0andQ2_0_g64kernels are compared against the shipped ggml shared objects of llama.cpp b11192. The check covers every weight, everyq8_0activation row and every dot product, and matches the compiler's fused multiply-add rounding. No other project was found that does this against the binary. The nearest are:- OxiLLaMa, which claims logit parity within a tolerance;
- Frink, whose quantizer output is byte-identical but whose matmul is not claimed to be;
- bitnet-rs, which has byte-level unit tests against its C++ baseline.
llama.cpp's own PR #26348 shows why the FMA detail matters: its kernel matched bit for bit in isolation, but perplexity still shifted once the compiler contracted it differently.
A plain-AVX2 kernel for
Q2_0_g64, a gap upstream has not closed. At b11192, llama.cpp has only scalar code forQ2_0on x86:- PR #24448 (merged 2026-07-07) is "CPU only, ARM NEON + generic scalar fallback";
- the x86 kernel in PR #26348 needs AVX-VNNI and is still open.
bankml's AVX2 kernel is 9.5β9.8Γ the reference per projection, bit-exact. Caveat: that speedup is measured against a scalar baseline. PrismML's own fork has AVX2 kernels for its other ternary format (
PQ2_0, PR #206, about 3Γ over its own baseline), which is a different format and not directly comparable.A pinned-model gateway with a receipt on every answer, in one zero-dependency static binary. Model hashes and answer hashes are common inside TEE and attestation stacks, but rare as a simple local gate. Caveat: bankml's receipts are not signed or hardware-attested. They prove what this gateway served, not where or by whom (Β§3 below).
Where others are clearly ahead
- A full forward pass in Rust. candle,
mistral.rs, Frink,
OxiLLaMa and bitnet-rs run
whole models natively. bankml answers through llama-server; its own Qwen3 forward pass is phase P3.
Since then: bankml's own forward pass generates llama.cpp's tokens since 0.2.7 (1-bit) and 0.2.8 (ternary), and
bankml serve --nativeanswers from it since 0.3.0; the Llama graph followed in 0.3.4 (TECHNICAL.md Β§IV.4). Unlike the engines above, it is token-identical to llama-server b11192 in the release gate (oracles.md). - GPUs. mistral.rs, Frink and PrismML's fork run on CUDA and/or Metal; bankml is CPU-only by design. Since then: a Vulkan worker with bankml's own SPIR-V takes a share of each 1-bit product, used only after it is bit-exact on the card (0.2.12β0.2.14). On the one integrated card measured (a Radeon Vega 3) decode did not change within noise (modules/gpu.md). The engines above remain far ahead on GPUs.
- Breadth of formats. OxiLLaMa lists K, IQ, TQ1_0, TQ2_0 and Q1_0_G128 (all "Alpha"); bankml has its own kernels for two formats and serves the rest through the reference. Since then: a third, F16 (0.3.4); every other format is still served only through the reference.
- ARM. Upstream
Q1_0/Q2_0have NEON; bankml's kernels are x86 AVX2 only. - End-to-end speed leadership. It is not shown: bankml's gains are per kernel, and at the whole-token level the decode-budget ratio against the reference ranges from 0.87Γ to 1.62Γ across thread counts in the 0.1.5 gate record, measured on a laptop that was also serving a model, so it is noisy and not a claim of end-to-end leadership. Since then: with its own forward pass bankml decodes the ternary model at 2.32β2.41 tokens/s against llama-server's 0.30 on the same laptop (CHANGELOG 0.2.8). On the 1-bit model llama-server was still faster at 0.3.0 (2.8 against 1.9β2.0 tokens/s), and parity there is not yet claimed (PERFORMANCE.md).
- Attestation. EigenAI (bit-exact GPU inference on a modified llama.cpp, re-executed in a TEE) and TEE serving stacks give receipts a third party can trust. bankml's receipts do not.
1. Rust-native inference engines
| project | what it runs on CPU | own kernels? | vs llama.cpp on CPU | exactness testing |
|---|---|---|---|---|
| candle (Hugging Face) | GGUF K-quants (k_quants.rs); x86 repack in progress, aarch64 in PR #3697 | yes (AVX2/NEON) | no published parity numbers | none found |
| mistral.rs | GGUF Q/K, ISQ, GPTQ, AWQ, HQQ, FP8 | yes (on candle) | published wins are CUDA; no CPU comparison | none found |
| burn / models | a general deep-learning framework | yes | not a GGUF-quant CPU engine | β |
| rustformers/llm | archived, "Unmaintained" (last push 2024-06) | called ggml | β | β |
| lm.rs | a minimal engine (last push 2024-10) | yes | β | β |
| ratchet | browser / WebGPU | β | β | β |
| Crane, kalosm | built on candle | β | β | β |
| OxiLLaMa (2026) | K, IQ, TQ1_0/TQ2_0, Q1_0_G128, all "Alpha" | yes (AVX2/AVX-512/NEON) | targets β₯ 80 % of llama.cpp | top-1 logit parity 32/32, within tolerance, not bit-exact |
| Frink (2026; blog) | K, IQ, MXFP4, MoE; CPU, Metal, CUDA | yes | 1.41β5.06Γ slower than llama.cpp on CPU, by its author | quantizer bytes identical on Q8_0, IQ4_NL |
| llama-gguf | GGUF | yes | 0.3 tok/s on Mistral-7B; "correctness over speed" | β |
2. 1-bit and ternary inference
Upstream llama.cpp (releases; latest b11243):
Q1_0arrived in PR #21273 (merged 2026-04-06, NEON and scalar only), and x86 SSSE3βAVX2+FMA in PR #21636 (merged 2026-04-20).Q2_0is in PR #24448 (merged 2026-07-07; x86 "later").- The only x86
Q2_0kernel is VNNI-only, in PR #26348 (open; 2.1β3.6Γ), which does not help a plain-AVX2 CPU such as bankml's test Ryzen 3200U. - The older ternary types
TQ1_0/TQ2_0come from PR #10010, and the formats are discussed in #22019. - Issue #29351 (2026-09-24) found an x86
Q8_0sign bug by checking against the generic C kernel, which is the method bankml's oracle uses.
PrismML (the Bonsai models): Bonsai-8B, announcement, docs, and the fork, which has x86 kernels for PrismML's newer formats:
PQ2_0AVX2/AVX-VNNI: PR #206, merged 2026-09-21, about 3Γ decode on Ternary-Bonsai-2-27B;Q1_04Γ8 repack: PR #233;PQ1_0: PR #255;PTQ1_0: PR #250.
Microsoft:
- bitnet.cpp has kernels I2_S, TL1 (2 weights β 4-bit index) and TL2 (3 weights β 5-bit). It reports 2.37β6.17Γ on x86 and 72β82 % less energy.
- T-MAC does lookup-table mixed-precision GEMM, up to 6.6Γ per kernel and 2.8Γ end to end over llama.cpp.
- Both work on BitNet-style ternary weights, a different format from ggml's
Q2_0.
In Rust:
- bitnet-rs (NEON);
- alice-aegis (2026), a
no_stdUEFI BitNet engine with frozen integer semantics, bit-identical digests across ISAs and SHA-256 receipts chaining the logits. It is the closest peer to bankml's receipts, but it checks against itself, not against llama.cpp; - bitnet-toy;
- bitnet-llm (FFI bindings).
No Rust engine other than bankml was found with Q2_0_g64; OxiLLaMa lists Q1_0_G128.
3. Verifiable and attested inference
| approach | projects | what it proves | bankml's relation |
|---|---|---|---|
| zero-knowledge proofs | EZKL (Rust, small models); zkLLM; DeepProve; OpenLLM | that a computation was done, with no trust in the prover; Hollow-LLM shows ZK correctness alone does not prove a real large model ran | none; far heavier |
| optimistic re-execution | opML (paper); Optimistic TEE-Rollups | fraud is challengeable by re-running | bankml's bit-exactness is what re-execution needs, but it has no dispute protocol |
| TEEs | EigenAI; OpenPCC; Phala private AI inference | signed receipts from attested hardware | bankml's receipts are unsigned, with no hardware root |
| model pin + answer hash | bankml | which verified file served, and that the text and request are unchanged | honest scope: integrity between a client and its own gateway |
4. What this means for bankml's roadmap
- P3 (its own forward pass) is where the field already is, so it is not optional if bankml is to stand beside candle and OxiLLaMa. Its distinct contribution would be a forward pass bit-exact against the compiled reference, which nobody else claims. Done since: 0.2.1β0.3.0 (BUILD_HISTORY.md).
- A plain-AVX2
Q2_0kernel is worth upstreaming (or offering to PrismML): the gap is real, current and documented upstream. Prepared since:upstream/q2_0_avx2.c(0.2.0), bit-exact against the shipped library on 200,000 of 200,000 random cases; submitting it is the authors' decision, and nothing has been submitted. - Receipts: signing them (an operator key) and binding them to a re-executable bit-exact run would move them from integrity toward verifiability, the direction EigenAI shows.
- NEON would bring the kernels to ARM, where upstream
Q1_0/Q2_0already are.
Papers
Binary and ternary networks
- Courbariaux, Bengio, David. BinaryConnect: Training Deep Neural Networks with Binary Weights during Propagations. NeurIPS 2015.
- Rastegari, Ordonez, Redmon, Farhadi. XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks. ECCV 2016.
- Li, Zhang, Liu. Ternary Weight Networks. 2016.
- Zhu, Han, Mao, Dally. Trained Ternary Quantization. ICLR 2017.
Quantized and 1-bit language models 5. Dettmers, Lewis, Belkada, Zettlemoyer. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale. NeurIPS 2022. 6. Frantar, Ashkboos, Hoefler, Alistarh. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. ICLR 2023. 7. Wang et al. BitNet: Scaling 1-bit Transformers for Large Language Models. 2023. 8. Ma et al. The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits. 2024. 9. Ma et al. BitNet b1.58 2B4T Technical Report. 2025.
CPU kernels for 1-bit and ternary inference 10. Wang, Zhou, Song et al. 1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs. 2024. 11. Wang, Zhou, Song et al. Bitnet.cpp: Efficient Edge Inference for Ternary LLMs. 2025. 12. Wei et al. T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge. EuroSys 2025. 13. Li, Yin, Wang et al. Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference. 2025.
Verifiable and attested inference 14. Sun, Li, Zhang. zkLLM: Zero Knowledge Proofs for Large Language Models. CCS 2024. 15. Conway, So, Yu et al. opML: Optimistic Machine Learning on Blockchain. 2024. 16. Chan, Ding, Chen et al. Optimistic TEE-Rollups. 2025. 17. Alves, Patankar, Pereira et al. EigenAI: Deterministic Inference, Verifiable Results. 2026. 18. Cankaya. Bit-Exact AI Inference Verification Without Performance Tradeoffs. 2026 (vLLM and Hugging Face on GPU, not llama.cpp). 19. Gailly et al. DeepProve: Verifiable End-to-End LLM Inference. IACR ePrint 2026/1112. 20. Yang, Ren, Xu, Zhang et al. OpenLLM: Modular and Scalable zkSNARKs for Verifiable LLM Inference. IACR ePrint 2026/1578. 21. Gong, Liu, Li. Hollow-LLM Attack. 2026. 22. Zhou, Zhao, Wang et al. OpenPCC: Open and Confidential LLM Serving on Commodity TEEs. 2026.
Retrieval, embedding and commitments used by bankml 23. Cormack, Clarke, BΓΌttcher. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. SIGIR 2009. doi:10.1145/1571941.1572114 (the DOI refuses scripts; the author's PDF resolves). 24. Chen, Xiao, Zhang et al. M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation (BGE-M3). 2024. 25. Laurie, Langley, Kasper. RFC 6962: Certificate Transparency. 2013. 26. Laurie, Messeri, Stradling. RFC 9162: Certificate Transparency Version 2.0. 2021.
Unverified details, stated as found:
- The author lists of papers 3, 4, 12, 25 and 26 are as commonly cited, not re-fetched.
- The venues of papers 12 and 14 come from secondary pages.
- Every project's own speed claims are the project's, not reproduced here.