|
Download docs/oracles.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 50.3 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/0ab4e0b6dd5a713718f7d8e1409d62164949e896/docs/oracles.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml@0ab4e0b6dd5a713718f7d8e1409d62164949e896/docs/oracles.md
-
curl -L -o oracles.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/0ab4e0b6dd5a713718f7d8e1409d62164949e896/docs/oracles.md
50.3 kB
| # Oracles in bankml | |
| ## What an oracle is here | |
| An **oracle** is an answer bankml did not compute and cannot influence, against which its own answer is compared. | |
| bankml's rule since 0.0.1 is that a result counts only if an oracle has checked it: *the same bits first, then the | |
| speed*. A faster kernel that is not bit-identical to its oracle is not a result; a construction (a hash, a Merkle | |
| root, an ABI encoding) that does not reproduce a published value is not trusted. | |
| The strongest oracle is the **compiled reference itself**: llama.cpp b11192's own shared libraries, called in-process | |
| through their exported symbols on the same bytes bankml reads. The others are published values: test vectors from | |
| standards (FIPS 180-4, RFC 6962), values published by the systems bankml interoperates with (Savante's doctrine root, | |
| the THOT spec's vectors, Hugging Face's and Ollama's sha256 for a model file), and independent implementations | |
| (pycryptodome, Foundry's `cast`, the Python guard bankml's Rust guard was ported from). | |
| This file lists every oracle bankml uses, what it checks, how to run it, and what it last found. The results are in | |
| [`testing/results/`](../testing/results/), one record per release, written by | |
| [`testing/release_gate.sh`](../testing/release_gate.sh). The latest record is 0.3.6's; figures marked 0.3.7, 0.3.8 or | |
| 0.3.9 are from [CHANGELOG.md](../CHANGELOG.md) for versions not yet released, and reach `testing/results/` when each | |
| is cut. Each module's own oracles are also listed on its page in [modules/](modules/README.md), under *How it is | |
| verified*. | |
| ## The release gate, stage by stage | |
| Every stage of `testing/release_gate.sh`, in the order it runs. A stage that fails stops the gate. The rows marked | |
| *regression* compare bankml with itself or with its own scalar model; they guard against change, and by the rule at the | |
| end of this file they are not oracles. | |
| **Always (no model needed)** | |
| | stage | checks | against | Β§ | | |
| |---|---|---|---| | |
| | `cargo test`, `cargo test -p bankml-capi` | unit tests, the scalar models of Β§2, the gateway's receipts against a mock server | scalar models of ggml; a mock llama-server (*regression* where it is bankml against bankml) | Β§2, Β§5 | | |
| | `cargo clippy -D warnings` | lint, both packages | β | β | | |
| | `spdx_check.py` | every tracked source file names its licence in an SPDX header matching its layer | LICENSING.md | β | | |
| | `test_gguf_guard.py` | the Python guard's own suite | its fixtures | Β§3 | | |
| | `test_ui.py` | the UI's data layer: keccak256, THOT manifests, CIDv1, the RFC 6962 tree | published values | Β§4 | | |
| | `test_console.py` (0.3.7) | the console's routes, loopback and CSP rules, the SELF block's units, the persona's doctrine root, the Infotags metadata | RFC 6962 roots; *regression* otherwise | Β§4 | | |
| | `test_connectors.py` | PostgreSQL publish and load on a throwaway cluster; a tampered row refused | the manifest's bytes | Β§4 | | |
| | `test_chain.py` | iNFT mint and load on a throwaway anvil devnet | the compiled contract | Β§4 | | |
| | `test_models.py` | the model importer: the pin, the licence gate, resume, a carrier started and rolled back | published sha256 values | Β§4 | | |
| | `guard_agree.py` | the Rust guard against the Python guard, JSON for JSON | the Python original | Β§3 | | |
| | `capi_oracle.py --printf` | `bankml_log` byte for byte | libc `snprintf` | Β§5b | | |
| **With the b11192 release (`BANKML_GGML_LIB`)**, each a Rust test run by name (`cargo test --release -- --ignored | |
| --exact`), then the live checks against a running `bankml serve --native`: | |
| | stage | checks | against | Β§ | | |
| |---|---|---|---| | |
| | `schema_oracle.py`, `content_oracle.py` (when `LLAMA_SRC` is set) | re-record the schema grammars and the content rule | llama.cpp's own code in `libllama-common.so` | Β§5e | | |
| | `sort_oracle.cpp` (when `g++` is present) | re-record `std::sort`'s orders | libstdc++ | Β§5f | | |
| | `oracle_tokenizer` | token ids, 4,346 cases | llama-server `/tokenize` | Β§1b | | |
| | `oracle_chat_template` | prompts byte for byte, 317 conversations | llama-server `/apply-template` | Β§1c | | |
| | `oracle_forward_embed_norm`, `oracle_forward_qkv_rope`, `oracle_forward_attention`, `oracle_forward_attention_tiled`, `oracle_forward_attention_split`, `oracle_forward_swiglu_sweep` | layer 0 operation by operation, the three attention kernels, ggml's `expf` | the shipped ggml | Β§1d | | |
| | `oracle_forward_model`, `oracle_forward_model_ternary` | every layer, `result_norm` and the logits | the shipped ggml graph | Β§1d | | |
| | `oracle_greedy_llama_server`, `_ternary`, `_long`, `_deep` | greedy tokens, short, long and deep prompts | llama-server | Β§1d | | |
| | `oracle_sample_llama_server` | seeded sampling, 40 continuations | llama-server | Β§1d | | |
| | `oracle_native_serve` | three conversations, 9 turns: text, counts, cache reuse | llama-server `/v1` | Β§1d | | |
| | `oracle_grammar_masks` | whole-vocabulary masks, both paths since 0.3.9 | libllama's grammar sampler | Β§5c, Β§5h | | |
| | `oracle_json_mode`, `oracle_json_mode_ternary` | answers under JSON mode and user grammars | llama-server | Β§5c | | |
| | `oracle_schema_grammars`, `oracle_json_schema`, `oracle_json_schema_ternary`, `oracle_json_schema_o4`, `oracle_json_content` | schema grammars, answers under schemas, the content rule | llama.cpp's own code; llama-server | Β§5e | | |
| | `oracle_persona_layer` | mindX's persona Modelfile, two ways | llama-server with the persona as system message | Β§5e | | |
| | `oracle_penalties`, `oracle_penalties_8b` | the repeat, frequency and presence penalties | llama-server | Β§5f | | |
| | `oracle_samplers`, `oracle_samplers_8b` (0.3.7) | typical-p, top-n-Ο, XTC, dynamic temperature, DRY | llama-server | Β§5f | | |
| | `oracle_std_sort` (0.3.7) | `std::sort`'s order of equal keys | libstdc++ | Β§5f | | |
| | `oracle_ggml_b11192_q8_0_kv_kernels` (0.3.9) | the `q8_0` KV cache's quantizer and dot product | the shipped ggml | Β§5h | | |
| | `gpu_q1_0_mat_vec_bit_exact` | the GPU kernels on every usable card | the CPU kernel, itself proven against ggml | Β§1d (0.2.13) | | |
| | `oracle_train_script`, `oracle_train_imprint` | mindXtrain's author and score stages | mindXtrain's Python | Β§1d (0.2.13) | | |
| | `oracle_ggml_b11192_real_bonsai_1_7b`, `_real_bonsai_8b_q1_0`, `_real_ternary_bonsai_8b` | every weight, `q8_0` rows, dot products | the shipped ggml | Β§1 | | |
| | `oracle_ggml_b11192_f16` | F16 products, both of ggml's paths | the shipped ggml | Β§5d | | |
| | `oracle_tokenizer_smollm`, `oracle_chat_template_chatml`, `oracle_forward_model_bonsai_1_7b`, `oracle_forward_model_llama_f16`, `oracle_llama_server_bonsai_1_7b`, `oracle_llama_server_llama_f16`, `oracle_native_serve_o4` | the O4 models: tokenizer, templates, whole model, llama-server's tokens, conversations and JSON mode | llama-server; the shipped ggml | Β§5d | | |
| | `ab_vs_ggml`, `ab_vs_ggml_q2_0`, `bench_q1_0_prefill_act`, `bench_memory_floor`, `decode_budget_q1_0`, `decode_budget_q2_0` | speed: kernel A/Bs (each first checks agreement), the memory floor, one token's matmul budget | the shipped ggml's own kernels; measurements, not oracles | Β§1, [PERFORMANCE.md](PERFORMANCE.md) | | |
| | `serve_oracle_ollama_shape` | the conversations through `/api/chat`, `/v1`, and `/v1` after a reload | llama-server's record | Β§1d (0.3.1) | | |
| | `o4_live` | the conversations, JSON mode and schemas live, on Bonsai-1.7B, SmolLM2-135M-Instruct and `mindx-gen39` | llama-server's records | Β§5d, Β§5e | | |
| | `persona_oracle_live` | the created `mindx-gen39` through `/api/chat`, `/api/generate`, `/v1`; `/api/ps` | llama-server's record | Β§5e | | |
| | `penalty_oracle_live`, `sampler_oracle_live` | the penalties and the sampler chain through `/v1` and `/api/chat` | llama-server's records | Β§5f | | |
| | `context_oracle_live` (0.3.8) | the context limit | llama-server's record | Β§5g | | |
| | `slot_oracle_live` (0.3.8) | slot save, restore and erase | the engine's own empty-slot answer; llama-server's refusals | Β§5g | | |
| | `session_oracle_live` (0.3.8) | interleaved conversations; simultaneous requests | llama-server's record; the queue against itself | Β§5g | | |
| | `kv_oracle_live` (0.3.9) | answers over a `q8_0` KV cache | llama-server with `--cache-type-k/v q8_0` | Β§5h | | |
| | `logprobs_oracle_live` (0.3.8) | logprobs, streamed and not | llama-server's record | Β§5g | | |
| | `json_oracle_live`, `json_schema_oracle_live` | JSON mode and schemas live on the 8B model | llama-server's records | Β§5c, Β§5e | | |
| | `capi_chat_oracle` | `bankml_chat` | `serve --native`; llama-server's record | Β§5b | | |
| **When their inputs are present** | |
| | stage | checks | against | Β§ | | |
| |---|---|---|---| | |
| | `oracle_convert_b11192` (per directory under `.models/convert/`) | `bankml convert`'s GGUF, sha256 for sha256 | b11192's `convert_hf_to_gguf.py` | Β§5e | | |
| | `oracle_name_heuristics` (when `BANKML_LLAMA_SRC` is set) | model-name heuristics, 168 ids | gguf-py | Β§5e | | |
| ## 1. The ggml oracle: kernels bit-exact against the compiled llama.cpp | |
| ### What it checks | |
| For each real model, three things, every one to the bit: | |
| | # | quantity | bankml side | ggml side (llama.cpp b11192, shipped binaries) | | |
| |---|---|---|---| | |
| | 1 | every weight of every low-bit tensor, dequantized to f32 | `q1_0::dequantize_row`, `q2_0::dequantize_row` | `dequantize_row_q1_0` / `dequantize_row_q2_0` in `libggml-base.so` | | |
| | 2 | activation rows quantized to `q8_0` (32 signed bytes and a half-precision scale per block) | `quantize_row_q8_0` | `quantize_row_q8_0` in `libggml-cpu-haswell.so` | | |
| | 3 | dot products of a weight row with a `q8_0` row | `vec_dot_ref` (scalar model), `vec_dot` / `vec_dot_act` (AVX2), `vec_dot_act_scalar` | `ggml_vec_dot_q1_0_q8_0` (AVX2 path), `ggml_vec_dot_q1_0_q8_0_generic`, `ggml_vec_dot_q2_0_q8_0` | | |
| Dequantized tensors are compared by the sha256 of their little-endian f32 output (a whole 8-billion-weight model, | |
| tensor by tensor, without holding it in memory); quantized rows byte for byte; dot products by their f32 bit | |
| patterns (`to_bits()`), not within a tolerance. | |
| ### Why "the compiled library" and not "the source" | |
| The first port of ggml's generic C, done faithfully from the source, disagreed with the shipped library in the last | |
| bit. Reading the disassembly showed why: the compiler had fused a multiply and an add into one FMA instruction, which | |
| rounds once instead of twice. A source-level port cannot see that; the binary oracle catches it. bankml's kernels | |
| reproduce the float order and the fused multiply-adds of the **shipped** `libggml-cpu-haswell.so` (the variant ggml's | |
| backend scorer loads on an AVX2 CPU without AVX-512), and the oracle proves it on real tensors. | |
| It also shows the oracle can tell float orders apart. For the ternary format ggml ships two builds of the same generic | |
| C: the haswell variant (FMA) and the baseline x64 variant (SSE, no FMA). They disagree with each other on 19 of the | |
| 762 recorded ternary dot products. bankml carries a model of each (`vec_dot_ref` for haswell, `vec_dot_ref_nofma` for | |
| x64) and matches each one in 762 of 762. | |
| ### How it runs | |
| 1. **Record** (once per model): [`testing/ggml_oracle.py`](../testing/ggml_oracle.py) loads the release's | |
| `libggml-base.so` and `libggml-cpu-haswell.so` (and, for Q2_0, `libggml-cpu-x64.so`) with `ctypes`, calls | |
| `ggml_cpu_init()`, memory-maps the GGUF and writes: | |
| - `dequant.tsv`: tensor name Β· element count Β· sha256 of ggml's f32 output, for every tensor of the type; | |
| - `vecdot.bin`: for sampled rows of every tensor (3 per tensor by default): the f32 activation `x`, ggml's `q8_0` | |
| of it, and ggml's dot products (AVX2 and generic). | |
| The release tarball is checked by sha256 before use (`llama-b11192-bin-ubuntu-x64.tar.gz`, `34cf6fa5β¦81ec7`). | |
| 2. **Compare** (every gate): the Rust tests `oracle_ggml_b11192_real_*` (in [`q1_0.rs`](../bankML/q1_0.rs) and | |
| [`q2_0.rs`](../bankML/q2_0.rs), `#[ignore]`d because they need the models) re-derive every recorded quantity from the same | |
| file through bankml's own memory map and assert equality. | |
| 3. **A/B against the live library** (every gate): `ab_vs_ggml` and `ab_vs_ggml_q2_0` `dlopen` the shipped haswell | |
| library directly (no crate, `dlopen`/`dlsym` declared by hand) and time ggml's kernel and bankml's on the same | |
| real weights in one process, after first checking they agree. | |
| ### What it found (0.1.7β0.2.0 gates; identical in each, and again under Rust 1.99 in 0.3.2) | |
| The gate runs these as `oracle_ggml_b11192_real_bonsai_1_7b`, `oracle_ggml_b11192_real_bonsai_8b_q1_0` and | |
| `oracle_ggml_b11192_real_ternary_bonsai_8b`. | |
| | model | tensors | weights | q8_0 rows | dot products | | |
| |---|---|---|---|---| | |
| | Bonsai-1.7B, `Q1_0` | 197 | 1,719,904,256 | 788 byte-exact | 788 bit-exact (AVX2 and `vec_dot_act`); generic port == ggml generic 788/788 | | |
| | Bonsai-8B, `Q1_0` | 254 | 8,188,239,872 | 762 byte-exact | 762 bit-exact; generic 762/762 | | |
| | Ternary-Bonsai-8B, `Q2_0_g64` | 254 | 8,188,239,872 | 762 byte-exact | 762 bit-exact vs haswell (reference, AVX2 and scalar paths); no-FMA model == x64 build 762/762 | | |
| ### Run it | |
| ```sh | |
| # once: the oracle files (numpy + the b11192 release) | |
| python3 testing/ggml_oracle.py .models/Bonsai-1.7B-Q1_0.gguf /path/to/llama-b11192 .models/oracle | |
| python3 testing/ggml_oracle.py .models/Bonsai-8B-Q1_0.gguf /path/to/llama-b11192 .models/oracle-8b-q1 | |
| python3 testing/ggml_oracle.py .models/Ternary-Bonsai-8B-Q2_0_g64.gguf /path/to/llama-b11192 .models/oracle-ternary | |
| # every time | |
| cargo test --release -- --ignored oracle_ggml_b11192 --nocapture --test-threads=1 | |
| BANKML_GGML_LIB=/path/to/llama-b11192 cargo test --release -- --ignored ab_vs_ggml --nocapture --test-threads=1 | |
| ``` | |
| ## 1b. The tokenizer oracle: token-identical to llama.cpp (0.2.1) | |
| P3, bankml's own forward pass, starts with the tokenizer. [`testing/tokenizer_oracle.py`](../testing/tokenizer_oracle.py) | |
| asks a running llama-server (b11192, the Bonsai / Qwen3 vocabulary) to tokenize a corpus, with special tokens parsed | |
| and not, and records every answer. The corpus is: | |
| - every document in the repository and Savante's canon texts; | |
| - the chat template's markers; | |
| - hand-picked edge cases: contractions, CRLF, every kind of whitespace, digits in several scripts, combining marks, | |
| CJK, right-to-left scripts, emoji with joiners, mathematical alphanumerics; | |
| - a seeded fuzz set of 2,000 strings drawn from 17 Unicode ranges. | |
| `tokenizer::tests::oracle_tokenizer` re-derives every case with bankml's tokenizer (`tokenizer.rs`, no crates) and | |
| requires the same ids in the same order. Last result: **4,346 of 4,346 cases token-identical** (the corpus grew at 0.3.4; 4,258 at 0.2.1), on the Qwen2 and the SmolLM2 pre-tokenizer each. | |
| Two details the oracle settles: | |
| - **Which special tokens split the text.** Qwen3's `<think>` markers are USER_DEFINED and split the text even when | |
| special tokens are not parsed; CONTROL tokens such as `<|im_start|>` do not. | |
| - **The pre-tokenizer's letter class.** It is Unicode general category L, which is not `char::is_alphabetic`. It is | |
| generated into `unicode_letters.rs`. | |
| ## 1c. The chat-template oracle: byte-identical prompts (0.2.2) | |
| A conversation becomes a prompt through the model's chat template, a Jinja program inside the GGUF. bankml does not | |
| run Jinja. `chat.rs` writes the Bonsai / Qwen3 template's rules out, and `check_template` accepts only that template, | |
| identified by the sha256 of its text. [`testing/template_oracle.py`](../testing/template_oracle.py) records | |
| llama-server's own `/apply-template` for 317 conversations. They cover: | |
| - a system prompt first, later, or absent; | |
| - assistant turns with and without `<think>` blocks, before and after the last real user query; | |
| - `reasoning_content`; | |
| - runs of tool results; | |
| - user messages that look like tool responses; | |
| - special markers and Unicode inside content; | |
| - 300 random conversations. | |
| `oracle_chat_template` requires every prompt **byte-identical**: **317 of 317**. The oracle found one server behaviour | |
| that the template alone would not predict: an empty `reasoning_content` is dropped before templating. | |
| ## 1d. The forward-pass oracle: llama.cpp's own operations, from the shipped ggml (0.2.3) | |
| The release has no tool that prints a model's intermediate values, so the oracle builds them the way the kernel oracle | |
| does. [`testing/forward_oracle.py`](../testing/forward_oracle.py) drives the **shipped** `libggml` through ctypes with | |
| the same operations llama.cpp's Qwen3 graph uses and computes them with the release's own CPU backend. For 300 real | |
| token ids (the chat markers, the table's edges, and tokens of the tokenizer's corpus) it records the sha256 of each | |
| row's f32 bytes: | |
| - `get_rows` on the Q1_0 token embedding: llama.cpp's `inp_embd`; | |
| - `rms_norm` with the model's epsilon, then `mul` by `blk.0.attn_norm.weight`: llama.cpp's `attn_norm-0`. | |
| bankml's `forward.rs` computes the same in the float order read from b11192's `ops.cpp`: | |
| - the sum of squares accumulated in double, one float product at a time, in order; | |
| - the mean rounded to float; | |
| - `scale = 1 / sqrtf(mean + eps)`; | |
| - each output `(x Β· scale) Β· w`, with no FMA. | |
| `oracle_forward_embed_norm` requires **300 of 300 rows bit-exact for both.** | |
| **Step four (0.2.4): Q, K, V, the head norms and RoPE.** The oracle also builds layer 0's attention inputs on a | |
| 28-token chat prompt: | |
| - `mul_mat` of the Q1_0 `attn_q`, `attn_k` and `attn_v` weights with the normed rows; | |
| - `reshape` to 128-wide heads, `rms_norm` and `mul` by `attn_q_norm` and `attn_k_norm`; | |
| - `rope_ext` in NEOX mode with the YaRN parameters llama.cpp's context derives (`freq_scale` 0.25, `ext_factor` 1, | |
| `attn_factor` 1.0, whose bits the record carries, beta 32 and 1, `n_ctx_orig` 16,384, base 10βΆ). | |
| RoPE runs twice: at positions 0β27, and at positions 7 to 63,214, where theta is a long product. | |
| `oracle_forward_qkv_rope` requires **140 of 140 rows bit-exact**. | |
| This step showed why the oracle is the shipped binary and not the source. Written as `ops.cpp` reads, bankml's RoPE | |
| matched only 30 of 84 rows, while the projections and norms before it were exact. The disassembly of | |
| `libggml-cpu-haswell.so` shows GCC's FMA contraction of the rotation, the YaRN mix and the magnitude term. With those | |
| three written as the same `mul_add`s, every row matches. | |
| **Step five (0.2.5): attention.** The oracle stores K and V through `cpy` to f16, as llama.cpp's cache does. It | |
| builds the causal f16 mask and runs `flash_attn_ext`, which is llama-server's default attention on the CPU, with F32 | |
| accumulation set as llama-graph sets it. The output then goes through `wo` (`mul_mat`) and the residual `add`. | |
| `oracle_forward_attention` requires **84 of 84 rows bit-exact** (`kqv_out`, `attn_out`, `ffn_inp`, 28 tokens). | |
| The reproduction follows ggml's reference path: | |
| - the f16 dot's four 8-lane FMA accumulators and their reduction tree; | |
| - an online softmax whose V accumulator is rounded to f16 at every step. | |
| The oracle rejects a near miss: with the softmax sum written as an FMA, only 59 of 84 rows match. At 0.2.5 two other | |
| kernels were not yet covered; each got its own oracle in 0.2.9 and 0.2.10 (below): | |
| - the tiled kernel, for 64 or more query rows; | |
| - the split-KV kernel, for a decode whose padded KV length reaches 512 (llama.cpp pads to multiples of 256, so from | |
| 257 cells in use). | |
| **Step six (0.2.6): the feed-forward block; layer 0 complete.** The oracle continues from `ffn_inp`: | |
| - `rms_norm` and `mul` by `ffn_norm`; | |
| - `mul_mat` by `ffn_gate` and `ffn_up`; | |
| - `swiglu_split`; | |
| - `mul_mat` by `ffn_down`; | |
| - `add`, giving `l_out`, the layer's output. | |
| `oracle_forward_attention` requires 112 of 112 rows through `l_out`. SiLU in the shipped build uses ggml's own | |
| vectorized `expf`, a polynomial in FMAs, not libm. bankml reproduces it lane by lane. | |
| A second check, `oracle_forward_swiglu_sweep`, feeds 24,600 values across Β±120 and the edges straight to | |
| `ggml_swiglu_split` and requires every output's bits (**24,600 of 24,600**). The sweep exists because the layer check | |
| alone was not enough: with libm's `expf`, `l_out` still matched 27 of 28 rows, since q8_0 quantization absorbs most | |
| of the difference, while the sweep matched only 19,245 of 24,600. | |
| **Step seven (0.2.7): the whole model, and llama-server itself.** | |
| - [`testing/model_oracle.py`](../testing/model_oracle.py) computes the whole Qwen3 graph in the shipped ggml: every | |
| layer, `output_norm` and `mul_mat` by `output.weight`. It runs one layer per ggml context and carries the residual | |
| stream between contexts as f32 bytes, which is the same arithmetic as one graph in a fraction of the memory. | |
| `oracle_forward_model` requires **1,064 of 1,064 rows bit-exact**: 36 Γ 28 `l_out` rows, 28 `result_norm` rows, | |
| and 28 logit rows of 151,669 values each. | |
| - [`testing/greedy_oracle.py`](../testing/greedy_oracle.py) steps outside the graph. It asks the running | |
| llama-server to render 6 chat prompts (`/apply-template`), tokenize them (`/tokenize`) and continue them greedily | |
| (`/completion`, top-k 1, prompt cache off). `oracle_greedy_llama_server` requires bankml's own forward pass to | |
| produce the **same tokens on every prompt (6 of 6, 164 tokens)**, and to end the turn where the server did. | |
| Both stayed inside the range bankml reproduced at 0.2.7: prompts under 64 tokens, contexts up to 256 cells (the | |
| largest case uses 80). Steps nine and ten removed both limits. llama.cpp pads the KV length to multiples of 256, so a decode beyond 256 cells takes the split-KV kernel. | |
| **Step eight (0.2.8): the ternary model.** Both whole-model checks run again on Ternary-Bonsai-8B (Q2_0_g64), whose | |
| every matrix, including the embedding and the output, is ternary: | |
| - `oracle_forward_model_ternary` checks bankml against the shipped ggml graph: **1,064 of 1,064 rows**. | |
| - `oracle_greedy_llama_server_ternary` checks against llama-server running the ternary model, with Savante's flags, | |
| on a spare port: **6 of 6 prompts, 140 tokens**. | |
| The greedy comparison now covers exactly the tokens the server generated. Its list includes the end-of-turn token | |
| when it produced one. Where it stopped otherwise, that is its stopping policy, not a token choice. | |
| **Step nine (0.2.9): long prompts and ggml's tiled kernel.** llama.cpp computes a prompt in micro-batches of up to | |
| 512 tokens. A micro-batch of 64 rows or more takes `flash_attn_ext_tiled`, a different algorithm from the reference | |
| path: f32 Q, a SIMD GEMM per 64-cell KV tile, a vectorized softmax summed in double, and an f32 accumulator. | |
| - The forward oracle runs layer 0 on a 150-row micro-batch in the shipped ggml. `oracle_forward_attention_tiled` | |
| requires **150 of 150 rows bit-exact**. The reference kernel would match only 1, the first row, so the kernel | |
| choice is itself checked. | |
| - `greedy_oracle.py --long` records llama-server's continuations for 6 prompts of 111β116 tokens. | |
| `oracle_greedy_llama_server_long` requires **the same tokens on all 6**. | |
| **Step ten (0.2.10): long contexts and ggml's split-KV kernel.** A single-token decode over 512 or more padded cells | |
| cuts the cells into one chunk per llama.cpp thread and reduces the partial softmaxes, so the thread count shapes the | |
| bits. | |
| - The forward oracle records one decode row at 7 context lengths (257β1,000 cells) and at 3 and 4 threads. | |
| `oracle_forward_attention_split` requires **14 of 14**. | |
| - `greedy_oracle.py --deep` records 3 continuations of 200 tokens that run to about 300 cells. | |
| `oracle_greedy_llama_server_deep` requires **the same 600 tokens**. | |
| The 4-thread cases did double duty. Besides showing the dependence on the thread count, they exposed the reduction's | |
| FMA contraction: at 3 threads the first chunk always held the maximum, where the fused and plain forms agree. | |
| **Step eleven (0.2.11): sampling.** `testing/sample_oracle.py` has llama-server sample 40 continuations with fixed | |
| seeds, and keeps the parameters the server reports it ran (`generation_settings`). `oracle_sample_llama_server` | |
| replays them through bankml's forward pass and `sampler.rs` and requires **every token (40 of 40, 1,175 tokens)**. | |
| The cases range over temperature 0β1.5, top-k 5β128, top-p and min-p. Top-k is libstdc++'s `partial_sort`, ported | |
| line for line, because the order it leaves tied logits in decides which token a draw lands on. | |
| **0.2.13: the GPU and mindXtrain.** | |
| - `gpu_q1_0_mat_vec_bit_exact` runs both Q1_0 GPU kernels on every usable card against the CPU kernel, which is | |
| bit-exact against ggml: on the Vega 3, every row of every shape from 64Γ128 to 12288Γ4096. `bankml gpu --verify` | |
| runs the same check for a user, and bankml trusts a card only after it passes. | |
| - `oracle_train_script` checks bankml's author stage against mindXtrain's `scripts.py` on 84 scripts, byte for | |
| byte. | |
| - `oracle_train_imprint` checks the score stage against `score_imprint` on 3,000 randomized cases, value for | |
| value. | |
| **0.2.14: the GPU inside the forward pass.** With a verified card present, `Weights::open` gives it a share of every | |
| 1-bit matrix's rows, so every token oracle runs with the GPU working. The whole-model check passes at 1,064 of | |
| 1,064 rows with the Vega 3 on 26 % of every matrix. | |
| The on-card oracle learned the lesson of this release. It now includes layer-shaped data: scales and magnitudes | |
| that vary per block, so products are inexact in f32. On that data the card's unfused `Fma` differed by one ulp, | |
| where random data had let it pass. bankml now computes the FMA exactly (BoldoβMelquiond), and with the driver's | |
| `Fma` put back, `bankml gpu --verify` refuses the card. | |
| **0.3.0: whole conversations through the server.** `testing/serve_oracle.py` holds three Savante-style | |
| conversations, three turns each, with fresh seeds at temperature 0.3. It sends them to a fresh llama-server's | |
| `/v1/chat/completions`, the path Savante uses, so the server's prompt cache carries each conversation from turn to | |
| turn, and records every answer with its prompt count, completion count and the prompt tokens taken from the cache. | |
| `oracle_native_serve` replays the conversations through bankML's native engine in one process and requires all of | |
| it on **9 of 9 turns**: the same text, the same counts, the same cache reuse. The rule it reproduces is visible in | |
| the numbers. Turn 2 reuses 49 of turn 1's 53 prompt tokens, because the empty `<think>` block that ended turn 1's | |
| prompt is not rendered when that answer becomes history. A new conversation reuses the 38 tokens of the shared | |
| system prompt. | |
| **0.3.1: the Ollama shape.** `testing/serve_oracle.py --bankml` starts `bankml serve --native` with the 8B 1-bit model | |
| on spare ports and sends the same conversations three ways, each from an empty slot (the model unloaded, then loaded | |
| again, so each run starts as a fresh llama-server does): through Ollama's `/api/chat`, with the first conversation | |
| streamed as NDJSON and the options mapped from Ollama's names; through `/v1/chat/completions`; and through `/v1` once | |
| more after an unload in which `/v1` itself loads, and verifies, the startup model. Every turn must give the same text, | |
| prompt and completion counts and cache reuse on every path, equal to llama-server b11192's record. So the Ollama | |
| path inherits the conversation oracle, and a reload changes nothing. The gate runs it as `serve_oracle_ollama_shape`. | |
| Every step of the forward pass above was added to these oracles before it counted, and every later feature (Β§5bβΒ§5h) | |
| follows the same rule. | |
| ## 2. Scalar models: every fast path against its own reference | |
| Between the real-model oracle runs, the kernels are held to a scalar model of ggml, on synthetic inputs, in the | |
| ordinary `cargo test`: | |
| | check | where | cases | | |
| |---|---|---| | |
| | half-precision conversion, all 65,536 values, and against the CPU's F16C instruction | `q1_0.rs` | 65,536 | | |
| | `q8_0` rounding (round-half-to-even, scale = 127 / max\|x\|) | `q1_0.rs` | edge cases | | |
| | AVX2 `Q1_0` dot == scalar model of ggml | `q1_0.rs` | 3,500 | | |
| | AVX2 and portable `Q2_0` dot == ggml model | `q2_0.rs` | 3,300 | | |
| | matrixβvector, matrixβmatrix and threaded paths == per-pair dots | `q1_0.rs`, `q2_0.rs`, `par.rs` | every shape tested | | |
| | the 0.0.4 prefill tile == per-pair | `q1_0.rs` (`act_tile_bit_exact`) | every tile | | |
| | the generic port == f64 arithmetic on dequantized data (the mathematical definition) | `q1_0.rs` | random blocks | | |
| A speed-up is admitted only after these and Β§1 pass on the same build. Kernels that were faster but not exact, or | |
| exact but not faster beyond noise, are kept with their numbers in [`testing/experiments/`](../testing/experiments/) | |
| (0.0.5). | |
| ## 3. The guard: Rust against the Python it was ported from | |
| bankml's header guard (`gguf.rs`: play, refuse or need more bytes; the three low-bit traps; hostile headers) was | |
| ported from minaiml's Python `gguf_guard.py`, vendored in [`testing/gguf_guard.py`](../testing/gguf_guard.py) with its | |
| suite. [`testing/guard_agree.py`](../testing/guard_agree.py) runs **both** on every synthetic case and on every real | |
| model present and requires identical JSON. Last result: **28/28 agree**. | |
| ## 4. Cryptographic constructions against published values | |
| | construction | oracle | where | | |
| |---|---|---| | |
| | SHA-256 (the model pin) | FIPS 180-4 test vectors; since 0.1.8 the SHA-NI path must also equal the portable rounds on every length 0β1,000 (and 4 KiB, 64 KiB, split updates), and a real 1.16 GB model must hash to coreutils `sha256sum`'s value and its published pin | `sha256.rs` (`fips_vectors`, `hardware_path_equals_portable_on_every_length`) | | |
| | the model pin | the sha256 the publisher lists: the FORK.json of the PYTHAI fork, a Hugging Face repository's LFS sha256 at a fixed revision, or the Ollama registry's layer digest; every import is hashed as it streams and kept only if equal | `bankml.rs` (`pin`), `sAGI/models.py` | | |
| | keccak256 (pure Python) | pycryptodome's keccak on every input length 0β400 bytes; and Savante's **published doctrine root** `0x92fe83ebβ¦ae137d0`, reproduced from her persona | `sAGI/agents.py`, `testing/test_ui.py` | | |
| | THOT manifests (`sagi.thot_manifest/1`) | the spec's own test vectors (`THOT_MANIFEST.md` Β§5: savante@1fcca89, jaimla@8b57ccf, luvai@0c1eef7): bundle root, Merkle root, identity CID | `sAGI/thot.py`, `testing/test_ui.py` | | |
| | CIDv1 (raw, sha2-256, base32) | the CID of `"abc"` that mindX's `rage.py` and Savante's ledger construction give | `testing/test_ui.py` | | |
| | `.history` Merkle tree (RFC 6962) | the Certificate Transparency reference roots for 1, 3 and 8 leaves | `testing/test_ui.py` | | |
| | iNFT ABI encoding | the compiled `iNFT_7857` contract itself, deployed from its artifact on a throwaway anvil devnet: it must accept the calldata bankml encodes (simulate and mint), refuse what it should (missing role, a repeated content root), and read back exactly the values encoded | `sAGI/chain.py`, `testing/test_chain.py` | | |
| | Savante's canon | her own offline verifier `bind/savante_verify.py`: 12 of 12 commitments, doctrine root | the Verifier tab | | |
| ## 5. The gateway: receipts against the text | |
| `bankml serve`'s receipts are checked end to end against a mock llama-server whose answers are known | |
| ([`testing/cli.rs`](../testing/cli.rs)): the `response_sha256` equals the sha256 of the text the mock sent (streamed and | |
| not), `request_sha256` equals the sha256 of the request body, and a model file changed after verification produces a | |
| 503 and no receipt. In use, the UI recomputes every answer's sha256 and marks β or `β received!`. | |
| ## 5b. The C API (0.3.2): the library against libc, and against the server | |
| Two oracles hold the C API (`capi/`, [CAPI.md](CAPI.md)), both run from C programs compiled with the system `cc`. | |
| - **`bankml_log` against libc `snprintf`** (`testing/capi/printf_oracle.c`). `bankml_log` is a C-variadic function | |
| defined in Rust, and its formatter is bankML's own, so the oracle is the C library's `snprintf`. Each case passes | |
| the same format and arguments to both. The message delivered to the sink must equal snprintf's output byte for | |
| byte, length included. The cases cover: | |
| - integers at `INT_MIN`/`LLONG_MIN`/`SIZE_MAX`, in every length; | |
| - `%f` ties (`%.0f` of 0.5, 1.5 and 2.5), `%.1100f` of the smallest subnormal, `DBL_MAX`, and `inf`/`nan` with | |
| signs; | |
| - strings, NULL, and an unterminated buffer bounded by a precision; | |
| - `%c` of 0 and 255, `%p`, `%%`, widths, precisions and `*`. | |
| Then the cases it does not support must give their exact markers. No argument may be read that cannot be typed, | |
| and `%n` must not write. The same comparison runs in `cargo test -p bankml-capi`, from Rust. | |
| - **`bankml_chat` against `serve --native` and llama-server** (`testing/capi/capi_oracle.py --chat`, the gate's | |
| `capi_chat_oracle`). The | |
| conversation oracle of Β§1d, carried across the C boundary: | |
| - on Ternary-Bonsai-8B, the same request bytes go to a live `serve --native` and then to `bankml_chat` in a fresh | |
| process. They must give the same text, the same streamed pieces, the same counts and cache reuse, and the same | |
| receipt hashes; | |
| - on Bonsai-8B Q1_0, `bankml_chat` must give llama-server b11192's recorded turns; | |
| - Bonsai-1.7B and an unpinned file must be refused, with `serve`'s reasons. | |
| ## 5c. JSON mode and grammars (0.3.3): llama.cpp's grammar sampler, and llama-server's answers | |
| Two oracles hold O6's first cut (`bankML/grammar.rs`). Neither compares bankML with itself. | |
| - **The grammar engine against llama.cpp's own** (`testing/grammar_oracle.py` β `oracle_grammar_masks`). | |
| `testing/grammar_oracle.cpp` drives libllama b11192's public API (`llama_sampler_init_grammar`, `_apply`, | |
| `_accept`) on the pinned vocabulary, loaded vocab-only. The grammar functions inside llama.cpp are not exported, | |
| but the sampler that wraps them is, so the oracle is the code the server runs, not a model of it. | |
| - **The vocabulary as the grammar reads it.** Every token's piece (`llama_token_to_piece`, special tokens | |
| rendered) and whether it ends generation. bankML must agree on **151,669 of 151,669** tokens and on the end set | |
| of 6. That set includes `<|fim_pad|>`, `<|repo_name|>`, `<|file_sep|>` and token 128247 `</s>`, which bankML's | |
| engine had not counted as ends before. | |
| - **Masks.** The whole-vocabulary mask, before every token and after the last, is recorded as how many tokens are | |
| allowed and an FNV-1a hash of which. The runs span 12 grammars: | |
| - llama-server's JSON-mode grammar, with its generation-prompt prefill; | |
| - all 8 grammars in llama.cpp's `grammars/` (json, json_arr, arithmetic, c, chess, english, japanese, list); | |
| - three written for what those miss: token terminals `<think>`, `!<|im_end|>` and `<[151644]>`; `.`; `{m,n}` | |
| edges; comments and CRLF; astral and CJK ranges. | |
| The inputs are valid, invalid, partial, escaped and unicode JSON, whitespace past `space`'s 20-character and | |
| 2-newline limits, prose, fences and truncations. Each goes in twice: tokenized as the tokenizer does, and as one | |
| token per byte, which splits every multi-byte character across tokens and exercises the partial-UTF-8 path. | |
| The result: **196 of 196 runs; 1,645 of 1,645 masks and 116 of 116 rejection points identical**. | |
| - **llama-server's answers under a grammar** (`testing/json_oracle.py --record` β `oracle_json_mode`, | |
| `oracle_json_mode_ternary`). llama-server b11192 runs with Savante's flags and `--verbose`, so that each answer | |
| reports its tokens and the grammar it ran. | |
| - **What is sent.** `/v1/chat/completions` with `response_format: {"type": "json_object"}` on 7 prompts, greedy | |
| and seeded at temperatures 0.7, 1.0 and 1.3 (top-k 40). Some prompts invite JSON. Others tempt prose, a haiku or | |
| code, so that the grammar has to reject what the model wants: 152 of 860 tokens on the 1-bit model were redrawn | |
| under the mask. One is a nested object of 159 tokens, one is unicode, and two answers are cut by `max_tokens`. | |
| The 1-bit model also gets two user `grammar` requests. | |
| - **What is required.** Each case runs from an empty cache. bankML's engine must give the same token ids, the end | |
| token included (one answer ends on `<|file_sep|>`), and the same raw text, `content`, finish reason and counts. | |
| The grammar text and generation prompt the server reports must be bankML's constant and prefill. | |
| - **The result:** Bonsai-8B Q1_0 **23 of 23** (860 tokens); Ternary-Bonsai-8B **13 of 13** (534 tokens, 49 redrawn). | |
| - **The live check** (`testing/json_oracle.py --bankml`, the gate's `json_oracle_live`) sends a subset through a running | |
| `serve --native`, each from an empty slot: `/v1` with `response_format`, streamed once; `/v1` with `grammar`; and | |
| Ollama's `/api/chat` with `format: "json"`. Every answer must equal the record: **8 of 8**. | |
| ## 5d. O4 (0.3.4): F16 products, the Llama graph, tied embeddings β the same oracles, three new models | |
| Every oracle family the 8B models have, extended to Bonsai-1.7B (tied embeddings), SmolLM2-135M-Instruct and | |
| `mindx-gen39` (the Llama graph in F16), each against llama-server b11192 running that very file; plus one new kernel | |
| oracle. | |
| - **F16 products against the shipped library** (`testing/f16_oracle.py` β `oracle_ggml_b11192_f16`). ggml multiplies | |
| an F16 weight two ways: `ggml_vec_dot_f16` for one column, llamafile's tinyBLAS for two or more (when the weight has | |
| a multiple of 4 rows and the row a multiple of 8). The oracle builds `ggml_mul_mat` graphs through ctypes and has | |
| the shipped `libggml-cpu-haswell.so` compute them: SmolLM2's real matrices (layers 0 and 29, and the 49,152-row token | |
| table) by 1, 2, 3, 5 and 28 columns, and synthetic shapes that reach the tails (`k % 32 β 0`) and the fallbacks | |
| (`rows % 4 β 0`, `k % 8 β 0`). Every F16 tensor widened by `ggml_fp16_to_fp32_row` is hashed too. The result: | |
| **211 of 211 tensors; 552,268 of 552,268 elements** in 87 products (61 tinyBLAS, 26 `vec_dot`). The two paths give | |
| different bits for the same row and column (`the_two_reductions_differ`), so the oracle checks the choice as well. | |
| - **The whole model** (`testing/model_oracle.py`, now for tied embeddings, F16 and the Llama graph: | |
| `rope_ext` mode 0, no Q/K norms, `ext_factor` 0): every layer's `l_out`, `result_norm` and the logits. An F16 | |
| model is replayed as one micro-batch with every row output, because that is what the oracle's graph computes. | |
| Bonsai-1.7B **840 of 840** rows, SmolLM2-135M-Instruct and mindx-gen39 **800 of 800** each | |
| (`oracle_forward_model_bonsai_1_7b`, `oracle_forward_model_llama_f16`). | |
| - **The tokenizer** (`tokenizer_oracle.py 127.0.0.1:PORT .models/oracle-tokenizer-smollm` β `oracle_tokenizer_smollm`): the same corpus on | |
| SmolLM2's vocabulary, **4,346 of 4,346**, with two new edge strings: bytes SmolLM2 has no token for, and plane-4 | |
| characters. Re-recorded on the Qwen3 vocabulary, the grown corpus caught a bug: `</s>` (NORMAL in the file, CONTROL | |
| in llama.cpp by its name) was taken as text in 2 of 4,346 cases. Fixed; 4,346 of 4,346. | |
| - **The templates** (`template_oracle.py`, written per model as `cases-<stem>.jsonl` β `oracle_chat_template_chatml`): | |
| SmolLM2-Instruct's and mindx-gen39's ChatML, **317 of 317** each. | |
| - **llama-server's tokens** (greedy short, `--long`, `--deep`; seeded; conversations; JSON mode, each recorded from that | |
| model's server; `oracle_llama_server_bonsai_1_7b`, `oracle_llama_server_llama_f16`, `oracle_native_serve_o4`): SmolLM2 6, 6, 3 and 40 of 40; mindx-gen39 6, 6, 3 and 40 of 40; Bonsai-1.7B 6 and 40 of 40; | |
| conversations 9 of 9 and JSON mode 23 of 23 on each. On a ChatML template llama-server builds a different JSON | |
| grammar (no `<think>` in its root); the record shows it, and bankML carries it per template. | |
| - **Live and through the C API**: `serve_oracle.py --bankml STEM [NAME]` and `json_oracle.py --bankml STEM [NAME]` on | |
| each model (mindx-gen39 asked for as `mindx-gen39`), and `capi_oracle.py --chat` against each model's record. The | |
| gate runs the live pair, with `json_schema_oracle.py --bankml` since 0.3.5, as its `o4_live` stage. | |
| ## 5e. JSON schemas and `bankml create` (0.3.5): llama.cpp's converter, parser and converter script, and the persona | |
| - **The schema grammar** (`testing/schema_oracle.{cpp,py}` β `oracle_schema_grammars`): llama.cpp b11192's own | |
| `json_schema_to_grammar` and `common_chat_templates_apply`, called inside the release's `libllama-common.so` (no | |
| model), on 173 schemas β llama.cpp's 81 test cases, Pydantic-shaped schemas like mindX's, edge cases and every | |
| refusal path β with the template read from each model that carries it (Bonsai's Qwen3, SmolLM2-Instruct's ChatML, | |
| mindx-gen39's ChatML). **148 grammars byte-identical bare and on the chat path, on each template**; every refusal | |
| with llama.cpp's message (24, 20). The ChatML wrapping (no reasoning rules, another root) was read from this output. | |
| - **The answers** (`testing/json_schema_oracle.py --record STEM` β `oracle_json_schema`, `oracle_json_schema_ternary`, | |
| `oracle_json_schema_o4`): | |
| llama-server b11192 asked under schemas through all three request shapes, greedy and seeded, each from an empty | |
| cache, with answers cut by `max_tokens` inside an object, a top-level string and a number. Required: the same token | |
| ids (end token included), raw text, content, finish and counts, and the server's grammar and generation prompt are | |
| bankML's. **Bonsai-8B 28 of 28, Ternary-Bonsai-8B 11 of 11, Bonsai-1.7B, SmolLM2-135M-Instruct and mindx-gen39 | |
| 56 of 56 each.** Live (`--bankml STEM [NAME]`, the gate's `json_schema_oracle_live` and `o4_live`): `/v1` (one | |
| streamed) and `/api/chat` `format: <schema>`. | |
| - **The content rule** (`testing/content_oracle.{cpp,py}` β `oracle_json_content`): llama.cpp's `common_chat_parse` | |
| itself, with the parser llama-server builds for each request, on every prefix of every recorded constrained answer | |
| and on edge cases: **30,063 of 30,063 texts on each template** (36 texts with a reasoning block, which no grammar | |
| admits after its prefill, are refused by llama.cpp's parser and listed). It found two differences the answer | |
| oracles had not: unfinished escapes, and llama-server's raw-text answer for an empty parse; both fixed. | |
| - **The converter** (`testing/convert_oracle.py`, `oracle_convert_b11192`, `oracle_name_heuristics`): bankml convert | |
| against llama.cpp b11192's `convert_hf_to_gguf.py --outtype f16` on the same directory, sha256 for sha256 β gen39 | |
| `6b64c748β¦` and SmolLM2 `ec30a679β¦`, run directly in 0.3.5 (torch 2.14.1+cpu) β and gguf-py's name heuristics on | |
| 168 of 168 ids. | |
| - **mindX's persona, end to end** (`testing/persona_oracle.py` β `oracle_persona_layer`, and live as | |
| `persona_oracle_live`): promote.py's | |
| persona Modelfile created two ways (`FROM` the merged directory; `FROM mindx-gen39` in place), the user's turns | |
| alone, against llama-server given the persona as the system message on the same GGUF: **27 of 27** token-identical | |
| for each; live through `/api/chat`, `/api/generate` and `/v1` from an empty slot each, and `/api/ps` names the | |
| derived model. (Without the empty slot, seeded answers moved: a cached prefix changes F16's product paths, so the | |
| live check starts each answer as the record did.) | |
| ## 5f. The penalties (0.3.6) and the rest of the sampler chain (0.3.7, unreleased) | |
| - **The penalties** (`testing/penalty_oracle.py` β `oracle_penalties`, `oracle_penalties_8b`). llama-server b11192 | |
| runs with Savante's flags from an empty cache, greedy and seeded, on 17 variants Γ 4 prompts made to repeat; the | |
| window reaches into the prompt, as llama-server fills it. Each case keeps the parameters the server read back. bankML | |
| must give the same tokens, and refuse what the server refuses with its message. **mindx-gen39 56 of 56 (2,478 | |
| tokens), Bonsai-1.7B 56 of 56 (1,895), Bonsai-8B 56 of 56 (1,568); 12 of 12 refusals each.** | |
| - **The rest of the default chain** (`penalty_oracle.py --kind sampler` β `oracle_samplers`, `oracle_samplers_8b`): | |
| typical-p, top-n-Ο, XTC, dynamic temperature and DRY, the same way, on 23 variants Γ 4 prompts β each sampler alone, | |
| typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and a disabling threshold, dynamic | |
| temperature at 0, DRY with defaults, custom breakers and the repeat penalty, all five at once. **mindx-gen39 76 of 76 | |
| (3,576 tokens), Bonsai-1.7B 76 of 76 (2,587); 16 of 16 refusals each.** The Bonsai-8B record is still to be taken | |
| (CHANGELOG 0.3.7). | |
| - **`std::sort`'s order** (`testing/sort_oracle.cpp` β `oracle_std_sort`). typical-p sorts with `std::sort`, which is | |
| not stable, so the order equal scores end in is the algorithm's own. The oracle is libstdc++ itself, compiled by the | |
| gate: **876 of 876 orders identical**, sizes 0 to 1,000, heavy ties, sorted, reversed and equal keys. | |
| - **Live** (`--bankml`, the gate's `penalty_oracle_live` and `sampler_oracle_live`): every recorded case through | |
| `/v1/chat/completions` and every fourth also through Ollama's `/api/chat`, each from an empty slot; a refusal must | |
| be a 400 with llama-server's message. The penalties: **85 of 85** on mindx-gen39. On `/api/chat` only typical-p of | |
| the 0.3.7 samplers exists; the others are not Ollama options and go through `/v1`. | |
| ## 5g. The serving contract (0.3.8, unreleased) | |
| Each records llama-server b11192's answers (`--record`) and replays the same requests through `bankml serve --native` | |
| on Bonsai-1.7B. | |
| - **The context limit** (`testing/context_oracle.py` β `context_oracle_live`). Both servers at a 256-token context, | |
| with context shift off (llama-server's default). Eight requests from a few tokens to past the context: the same text, | |
| counts and `finish_reason` (`"length"` at the full context), and for a prompt that does not fit the same 400 status | |
| and body (`exceed_context_size_error`). **8 of 8.** | |
| - **Slots** (`testing/slot_oracle.py` β `slot_oracle_live`). `POST /slots/0?action=save|restore|erase`. Here the | |
| answer to compare with is the engine's own: an answer after a restore, in the same server and after a restart, must | |
| equal the answer computed from an empty slot, and the restore must skip the prompt (`cache_n` > 0). By the rule at | |
| the end of this file that half is a regression test. The refusals are llama-server's (501 without a slot directory, | |
| "Invalid slot ID", "Invalid action", "Invalid filename", "Unable to restore slot: β¦"), and a file from another model | |
| or damaged by one bit is refused. **19 of 19.** | |
| - **One slot and the host prompt cache** (`testing/session_oracle.py` β `session_oracle_live`). Three conversations | |
| take turns (A1 B1 C1 A2 β¦) through llama-server with its prompt cache on, so each turn finds another conversation's | |
| tokens in the slot: bankML must give the same text, counts and `cache_n`, turn by turn | |
| ([modules/prompt_cache.md](modules/prompt_cache.md)). Then four requests sent at once must all be answered, none | |
| mixed, each with the text it gets alone; llama-server's order among simultaneous arrivals is not deterministic, so | |
| this half checks bankML against itself. **14 of 14.** | |
| - **Logprobs** (`testing/logprobs_oracle.py` β `logprobs_oracle_live`). `/v1` with `logprobs` and `top_logprobs`: the | |
| same token ids, texts and bytes, every logprob the same 32-bit float, llama-server's entry rules for stop words and | |
| split UTF-8, and its refusals. Five cases are streamed, and there every chunk's delta and entry must match. | |
| **14 of 14.** | |
| ## 5h. The q8_0 KV cache and the grammar trie (0.3.9, in progress) | |
| - **The KV cache's kernels** (`oracle_ggml_b11192_q8_0_kv_kernels`, `forward.rs`). The shipped haswell library, | |
| loaded with `dlopen`: `quantize_row_q8_0` byte for byte and `ggml_vec_dot_q8_0_q8_0` bit for bit, on random | |
| activations with outliers and on blocks of β128 and 127 where `maddubs` saturates. **4,000 rows quantized | |
| byte-exact, 4,000 dot products bit-exact.** | |
| - **Answers over the cache** (`testing/kv_oracle.py` β `kv_oracle_live`). llama-server b11192 with | |
| `--cache-type-k q8_0 --cache-type-v q8_0`, and bankML with `BANKML_CACHE_TYPE=q8_0`: greedy and seeded, a 2,244-token | |
| prompt, a 320-token answer and a two-turn conversation; text, counts and `cache_n`. **6 of 6 (567 tokens).** The | |
| first replay passed 2 of 6: llama.cpp rotates Q, K and V through a Hadamard transform around a quantized cache, which | |
| bankML now reproduces (`fwht`). | |
| - **The grammar mask through a trie** (`oracle_grammar_masks`, Β§5c). The oracle now computes every mask both ways, one | |
| candidate at a time and by the trie's walk, each against llama.cpp's mask: **196 of 196 runs, 1,645 of 1,645 masks** identical to llama.cpp b11192 by each. | |
| ## 6. Oracles planned | |
| - **P3, bankml's own forward pass: done.** The criterion fixed here before it was built β at temperature 0, on the | |
| same prompts, an answer **token-identical** to llama.cpp b11192's β is met by Β§1d's greedy oracles (0.2.7 on), | |
| under seeded sampling too (0.2.11), and for whole conversations (0.3.0). | |
| - **Planned** ([TODO.md](TODO.md)): tool calls, each with llama-server's answers; a 4-bit (`q4_0`) KV cache with the | |
| same rotation; `Q8_0` weights with an oracle for Qwen's own template; NEON kernels against llama.cpp's ARM build. | |
| - **Speculative decoding (measured in 0.1.8).** A draft model (Bonsai-1.7B) and n-gram speculation both left the | |
| output **token-identical** at temperature 0 on every run; neither was faster beyond this laptop's noise, so neither | |
| is the default (n-gram is an opt-in). The same criterion applies to any future speed-up that changes how tokens are | |
| computed. | |
| - **The upstream `Q2_0` kernel (0.2.0).** `upstream/test_q2_0_avx2.c` holds the C drop-in to the shipped | |
| `ggml_vec_dot_q2_0_q8_0`, bit for bit: 200,000 of 200,000 random cases. | |
| ## The rule, restated | |
| - An oracle is **external**: a compiled library, a standard's vectors, a published value, an independent | |
| implementation. A test that compares bankml with bankml is a regression test, not an oracle. | |
| - A comparison is **exact** where the quantity is exact: bits, bytes, hashes. Tolerances are used for nothing that an | |
| oracle can check exactly. | |
| - A **speed claim** is made only on code that passed every oracle in the same gate run, and only if the gain is beyond | |
| the run-to-run noise measured on the same machine. | |
| - Every gate run is **kept** (`testing/results/<version>.txt`), including the rejected experiments' numbers. | |