Gregory-L commited on
Commit
d1c88ba
·
verified ·
1 Parent(s): cc08b7a

Ask bankML: your own bankML (verified, receipt checked in the page) or a Hugging Face provider (labelled not bankML); source at github main cfe81e1 (serve --allow-origin)

Browse files
This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. CHANGELOG.md +11 -1
  2. LICENSING.md +1 -1
  3. README.md +12 -5
  4. bankML/main.rs +2 -1
  5. bankML/ollama.rs +1 -1
  6. bankML/serve.rs +51 -5
  7. bankml-chat.js +204 -0
  8. docs/BUILD_HISTORY.md +1 -1
  9. docs/index.html +4 -0
  10. docs/install.md +1 -0
  11. docs/modules/gguf.md +1 -1
  12. docs/thesis.md +106 -2
  13. docs/usage.md +3 -0
  14. index.html +59 -0
  15. sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg +0 -0
  16. sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg +0 -0
  17. sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg +0 -0
  18. sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg +0 -0
  19. sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg +0 -0
  20. sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg +0 -0
  21. sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg +0 -0
  22. sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg +0 -0
  23. sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg +0 -0
  24. sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg +0 -0
  25. sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg +0 -0
  26. sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg +0 -0
  27. sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg +0 -0
  28. sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg +0 -0
  29. sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg +0 -0
  30. sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg +0 -0
  31. sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg +0 -0
  32. sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg +0 -0
  33. sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg +0 -0
  34. sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg +0 -0
  35. sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg +0 -0
  36. sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg +0 -0
  37. sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg +0 -0
  38. sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg +0 -0
  39. sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg +0 -0
  40. sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg +0 -0
  41. sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg +0 -0
  42. sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg +0 -0
  43. sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg +0 -0
  44. sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg +0 -0
  45. sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg +0 -0
  46. sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg +0 -0
  47. sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg +0 -0
  48. sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg +0 -0
  49. sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg +0 -0
  50. sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg +0 -0
CHANGELOG.md CHANGED
@@ -27,6 +27,15 @@ decode speed, the third, is measured on an idle machine next.
27
  against 38.8 ms** (13×), p90 49.3 against 80.6 ms, over the oracle's 1,645 masks. The oracle now computes every
28
  mask both ways: **196 / 196 runs, 1,645 / 1,645 masks** identical to llama.cpp b11192 by each.
29
 
 
 
 
 
 
 
 
 
 
30
  ### The code audit and the documentation (2026-10-06)
31
  - **Comments, professional and short.** Every source file's comments were cut to the contract, the invariants, the
32
  safety reasoning and the exact upstream behaviour that keeps the bits; the history and rationale moved to its
@@ -155,7 +164,8 @@ measurements.
155
  × 4 prompts: each sampler alone, typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and
156
  a disabling threshold, dynamic temperature at 0, DRY with defaults, custom breakers and with the repeat penalty, all
157
  five at once): mindx-gen39 **76 / 76** answers token-identical (3,576 tokens), Bonsai-1.7B **76 / 76** (2,587
158
- tokens); **16 / 16** refusals each with llama-server's message. Bonsai-8B: to be recorded.
 
159
  - `oracle_std_sort` (`testing/sort_oracle.cpp`, libstdc++'s own `std::sort`): **876 / 876** orders identical — sizes 0
160
  to 1,000, heavy ties, sorted, reversed and equal keys — the order typical-p's unstable sort leaves equal scores in.
161
 
 
27
  against 38.8 ms** (13×), p90 49.3 against 80.6 ms, over the oracle's 1,645 masks. The oracle now computes every
28
  mask both ways: **196 / 196 runs, 1,645 / 1,645 masks** identical to llama.cpp b11192 by each.
29
 
30
+ ### Your own bankML from a web page (`--allow-origin`)
31
+ - `bankml serve … --allow-origin ORIGIN` lets one named web page call the gateway from a browser: its CORS preflight
32
+ is answered (with Chrome's private-network grant, as a public page reaching a loopback address requires) and its
33
+ answers, streamed or not, carry `Access-Control-Allow-Origin`; any other origin gets none and a 403 preflight. The
34
+ loopback `Host` and JSON-POST rules are unchanged. It is what lets the bankML Space's page
35
+ ([PYTHAI/bankml](https://huggingface.co/spaces/PYTHAI/bankml)) talk to a visitor's own bankML, free, on their own
36
+ CPU — the way Savante's page reaches a local engine. Checked live: the preflight, plain and streamed answers with
37
+ the receipt, and another origin refused.
38
+
39
  ### The code audit and the documentation (2026-10-06)
40
  - **Comments, professional and short.** Every source file's comments were cut to the contract, the invariants, the
41
  safety reasoning and the exact upstream behaviour that keeps the bits; the history and rationale moved to its
 
164
  × 4 prompts: each sampler alone, typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and
165
  a disabling threshold, dynamic temperature at 0, DRY with defaults, custom breakers and with the repeat penalty, all
166
  five at once): mindx-gen39 **76 / 76** answers token-identical (3,576 tokens), Bonsai-1.7B **76 / 76** (2,587
167
+ tokens), Bonsai-8B **76 / 76** (2,361 tokens, `oracle_samplers_8b`); **16 / 16** refusals on each, with
168
+ llama-server's message.
169
  - `oracle_std_sort` (`testing/sort_oracle.cpp`, libstdc++'s own `std::sort`): **876 / 876** orders identical — sizes 0
170
  to 1,000, heavy ties, sorted, reversed and equal keys — the order typical-p's unstable sort leaves equal scores in.
171
 
LICENSING.md CHANGED
@@ -34,7 +34,7 @@ Rules that keep the layers honest:
34
  | what | licence |
35
  |---|---|
36
  | `upstream/` (the AVX2 `Q2_0` kernel prepared for llama.cpp) | `MIT`, llama.cpp's licence, so it can be contributed as is |
37
- | `testing/gguf_guard.py` | vendored from minaiml (same authors); `MIT OR Apache-2.0` here |
38
  | `sAGI/voice/knobs/savante_knobs.js` | a build of DreamKnob and React (MIT); regenerate with `node sAGI/voice/knobs/build.mjs` |
39
  | `sAGI/voice/cache/`, `sAGI/voice/export/` | Savante's voice, rendered by bankml with Piper and the `en_GB-cori-high` voice (trained on public-domain LibriVox recordings); offered under `MIT OR Apache-2.0` |
40
  | models | not part of this repository; bankml imports only models with open-source licences and pins each by sha256 (`sAGI/models.py`) |
 
34
  | what | licence |
35
  |---|---|
36
  | `upstream/` (the AVX2 `Q2_0` kernel prepared for llama.cpp) | `MIT`, llama.cpp's licence, so it can be contributed as is |
37
+ | `testing/gguf_guard.py` | vendored from [minaiml](https://github.com/minaiml) (same authors); `MIT OR Apache-2.0` here |
38
  | `sAGI/voice/knobs/savante_knobs.js` | a build of DreamKnob and React (MIT); regenerate with `node sAGI/voice/knobs/build.mjs` |
39
  | `sAGI/voice/cache/`, `sAGI/voice/export/` | Savante's voice, rendered by bankml with Piper and the `en_GB-cori-high` voice (trained on public-domain LibriVox recordings); offered under `MIT OR Apache-2.0` |
40
  | models | not part of this repository; bankml imports only models with open-source licences and pins each by sha256 (`sAGI/models.py`) |
README.md CHANGED
@@ -5,6 +5,10 @@ colorFrom: green
5
  colorTo: indigo
6
  sdk: static
7
  app_file: index.html
 
 
 
 
8
  pinned: true
9
  license: mit
10
  short_description: Verified 1-bit and ternary LLMs on a CPU, in Rust
@@ -18,10 +22,12 @@ tags:
18
  - verified-inference
19
  ---
20
 
21
- > **This Space holds bankML's whole source**, and a page presenting it. It is ready to run the engine itself:
22
- > `Dockerfile` and `hf/start.sh` serve `bankml serve --native` with the ternary Bonsai-8B (verified against its sha256
23
- > pin before it can answer) and the bankML console in front, speaking as bankML (`sAGI/personas/bankml.persona`),
24
- > public and read-only. That needs Docker hardware on the Space (set `sdk: docker` and `app_port: 7860`). Source of record:
 
 
25
  > [github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml) · licence `MIT OR Apache-2.0`.
26
 
27
  <h1 align="center">bankML</h1>
@@ -48,7 +54,8 @@ tags:
48
  <p align="center">
49
  <a href="https://deltaverse.pythai.net/bankml"><b>Why bankML</b></a> (the short version, on the web) &middot;
50
  <a href="docs/thesis.md"><b>the thesis</b></a> &middot; <a href="docs/usage.md"><b>usage</b></a> &middot;
51
- <a href="CHANGELOG.md"><b>changelog</b></a>
 
52
  </p>
53
 
54
  ---
 
5
  colorTo: indigo
6
  sdk: static
7
  app_file: index.html
8
+ hf_oauth: true
9
+ hf_oauth_expiration_minutes: 480
10
+ hf_oauth_scopes:
11
+ - inference-api
12
  pinned: true
13
  license: mit
14
  short_description: Verified 1-bit and ternary LLMs on a CPU, in Rust
 
22
  - verified-inference
23
  ---
24
 
25
+ > **Ask bankML, two ways, never confused.** *Your own bankML*: this page talks to `bankml serve` on your machine
26
+ > (started with `--allow-origin https://pythai-bankml.static.hf.space`) — bankML's verified arithmetic on your CPU,
27
+ > free, with a receipt whose sha256 the page checks. *A Hugging Face provider*: sign in, and a provider-hosted model
28
+ > answers with bankML's persona (`sAGI/personas/bankml.persona`) on your inference quota — labelled on every answer
29
+ > as not bankML, with no receipt. This Space also holds bankML's whole source, and is ready to run the engine itself
30
+ > (`Dockerfile`, `hf/start.sh`) once it has Docker hardware. Source of record:
31
  > [github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml) · licence `MIT OR Apache-2.0`.
32
 
33
  <h1 align="center">bankML</h1>
 
54
  <p align="center">
55
  <a href="https://deltaverse.pythai.net/bankml"><b>Why bankML</b></a> (the short version, on the web) &middot;
56
  <a href="docs/thesis.md"><b>the thesis</b></a> &middot; <a href="docs/usage.md"><b>usage</b></a> &middot;
57
+ <a href="CHANGELOG.md"><b>changelog</b></a> &middot;
58
+ <a href="https://huggingface.co/spaces/PYTHAI/bankml"><b>on Hugging Face</b></a>
59
  </p>
60
 
61
  ---
bankML/main.rs CHANGED
@@ -13,7 +13,7 @@ const USAGE: &str = "usage: bankml usage [PID …]
13
  bankml pin FILE --fork FORK.json
14
  bankml verify FILE --fork FORK.json [--engine mainline|prism] [--json]
15
  bankml serve FILE --fork FORK.json [--upstream HOST:PORT | --spawn LLAMA_SERVER] [--listen HOST:PORT] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR]
16
- bankml serve FILE --fork FORK.json --native [--listen HOST:PORT] [--upstream HOST:PORT] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR]
17
  (answers from bankML's own forward pass; also serves the engine address;
18
  OpenAI /v1 and Ollama /api; --registry: every model pinned in DIR,
19
  default ~/.local/share/bankml/forks, by name, one resident at a time)
@@ -293,6 +293,7 @@ fn main() {
293
  // `--registry` without DIR: the importer's forks directory.
294
  registry: a.iter().any(|x| x == "--registry").then(|| registry_dir(&a)),
295
  keep_alive: opt("--keep-alive"),
 
296
  };
297
  match bankml::serve::run(cfg) {
298
  Ok(()) => 0,
 
13
  bankml pin FILE --fork FORK.json
14
  bankml verify FILE --fork FORK.json [--engine mainline|prism] [--json]
15
  bankml serve FILE --fork FORK.json [--upstream HOST:PORT | --spawn LLAMA_SERVER] [--listen HOST:PORT] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR]
16
+ bankml serve FILE --fork FORK.json --native [--listen HOST:PORT] [--upstream HOST:PORT] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR] [--allow-origin ORIGIN]
17
  (answers from bankML's own forward pass; also serves the engine address;
18
  OpenAI /v1 and Ollama /api; --registry: every model pinned in DIR,
19
  default ~/.local/share/bankml/forks, by name, one resident at a time)
 
293
  // `--registry` without DIR: the importer's forks directory.
294
  registry: a.iter().any(|x| x == "--registry").then(|| registry_dir(&a)),
295
  keep_alive: opt("--keep-alive"),
296
+ allow_origin: opt("--allow-origin"),
297
  };
298
  match bankml::serve::run(cfg) {
299
  Ok(()) => 0,
bankML/ollama.rs CHANGED
@@ -522,7 +522,7 @@ fn answer(c: &mut TcpStream, l: &Loaded, req: &Json, msgs: Option<&Json>, o: Opt
522
  };
523
  let mut content = crate::grammar::ContentStream::new(&constraint);
524
  if stream {
525
- write!(c, "HTTP/1.1 200 OK\r\nContent-Type: application/x-ndjson\r\nCache-Control: no-cache\r\nConnection: close\r\n\r\n")?;
526
  let send = |c: &mut TcpStream, piece: &str| piece.is_empty() || c.write_all(piece_line(model, chat, piece).as_bytes()).and_then(|_| c.flush()).is_ok();
527
  let done = eng.complete(&prompt, params, o.max, grammar, |piece| {
528
  if t.ttft.is_none() {
 
522
  };
523
  let mut content = crate::grammar::ContentStream::new(&constraint);
524
  if stream {
525
+ write!(c, "HTTP/1.1 200 OK\r\nContent-Type: application/x-ndjson\r\nCache-Control: no-cache\r\n{}Connection: close\r\n\r\n", crate::serve::cors_headers())?;
526
  let send = |c: &mut TcpStream, piece: &str| piece.is_empty() || c.write_all(piece_line(model, chat, piece).as_bytes()).and_then(|_| c.flush()).is_ok();
527
  let done = eng.complete(&prompt, params, o.max, grammar, |piece| {
528
  if t.ttft.is_none() {
bankML/serve.rs CHANGED
@@ -43,6 +43,25 @@ pub struct Config {
43
  pub registry: Option<PathBuf>,
44
  /// `--native`: how long `/api/*` keeps a model resident when the request does not say (Ollama's default, 5m)
45
  pub keep_alive: Option<String>,
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
46
  }
47
 
48
  struct State {
@@ -57,6 +76,8 @@ struct State {
57
  ident: FileIdent,
58
  /// where `/slots/{id}?action=save|restore` keeps slot files (`--slot-dir`; native mode)
59
  slot_dir: Option<PathBuf>,
 
 
60
  }
61
 
62
  /// The file identity checked before each P0 answer: device, inode, size and mtime, not the bytes (re-hashing per
@@ -110,6 +131,9 @@ mod sig {
110
 
111
  /// Verify the model, launch or check the upstream (or start `--native`), then serve until killed.
112
  pub fn run(cfg: Config) -> Result<(), String> {
 
 
 
113
  let before = ident(&cfg.model).map_err(|e| format!("{}: {e}", cfg.model.display()))?;
114
  let verified = crate::verify(&cfg.model, &cfg.fork_json, cfg.engine)?;
115
  let model = cfg.model.canonicalize().map_err(|e| format!("{}: {e}", cfg.model.display()))?;
@@ -153,7 +177,7 @@ pub fn run(cfg: Config) -> Result<(), String> {
153
  }
154
  let hashed_at = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
155
  let engine = "llama.cpp b11192 llama-server (loopback), behind bankml P0".to_string();
156
- let st = Arc::new(State { native: None, keep_alive: crate::native::KeepAlive::Forever, verified, model, upstream, engine, hashed_at, ident: id, slot_dir: None });
157
  let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
158
  eprintln!("bankml serve {}: {} verified (sha256 {}), upstream {} serves it; listening on http://{}",
159
  crate::VERSION, st.model.display(), st.verified.model_sha256, st.upstream, cfg.listen);
@@ -189,7 +213,7 @@ fn run_native(cfg: Config, verified: Verified, model: PathBuf, id: FileIdent, up
189
  if let Some(d) = &cfg.slot_dir {
190
  std::fs::create_dir_all(d).map_err(|e| format!("--slot-dir {}: {e}", d.display()))?;
191
  }
192
- let st = Arc::new(State { slot_dir: cfg.slot_dir.clone(), native: Some(rs), keep_alive: ka, verified: Verified { model_sha256: sha.clone(), guard: "play", engine: cfg.engine.as_str(), arch: None, name: None, types: Vec::new() },
193
  model, upstream: upstream.clone(), engine, hashed_at, ident: id });
194
  let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
195
  let lu = TcpListener::bind(&upstream).map_err(|e| format!("cannot listen on the engine address {upstream}: {e} (is llama-server running there?)"))?;
@@ -378,7 +402,7 @@ pub(crate) fn respond(c: &mut TcpStream, code: u16, ctype: &str, body: &[u8]) ->
378
  _ => "Error",
379
  };
380
  let code = if (100..600).contains(&code) { code } else { 502 }; // a garbled upstream head is a bad gateway
381
- write!(c, "HTTP/1.1 {code} {reason}\r\nContent-Type: {ctype}\r\nContent-Length: {}\r\nConnection: close\r\n\r\n", body.len())?;
382
  c.write_all(body)
383
  }
384
 
@@ -405,6 +429,18 @@ fn handle(mut c: TcpStream, st: &State) -> std::io::Result<()> {
405
  if !loopback_host(header(&h, "host").unwrap_or("")) {
406
  return refuse(403, b"bankml serve answers loopback clients only (Host must be 127.0.0.1, localhost or [::1])", &mut r);
407
  }
 
 
 
 
 
 
 
 
 
 
 
 
408
  if header(&h, "transfer-encoding").is_some() {
409
  return refuse(400, b"chunked requests are not accepted; send Content-Length", &mut r);
410
  }
@@ -583,7 +619,7 @@ fn native_chat(c: &mut TcpStream, rs: &crate::native::Residency, body: &[u8]) ->
583
  let model_id = l.model.file_name().map(|n| n.to_string_lossy().into_owned()).unwrap_or_default();
584
  let created = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
585
  let r = if stream {
586
- write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\nConnection: close\r\n\r\n")?;
587
  // with logprobs, llama-server's shape: the first token's partial opens with the role delta, and each token's
588
  // entry rides on the last delta its partial sends (no delta, no entry)
589
  let logprobs = nc.params.n_probs > 0;
@@ -978,7 +1014,7 @@ fn chat(c: &mut TcpStream, st: &State, body: &[u8]) -> std::io::Result<()> {
978
  let merged = format!("{}, \"bankml_receipt\": {}}}", &s[..end], t.receipt(&st.engine, &st.verified));
979
  return respond(c, 200, "application/json", merged.as_bytes());
980
  }
981
- write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\nConnection: close\r\n\r\n")?;
982
  let mut pending = Vec::new();
983
  while let Some(ch) = b.next(&mut r)? {
984
  pending.extend(ch);
@@ -1268,6 +1304,16 @@ mod tests {
1268
  assert_eq!(header(&h, "host"), Some("127.0.0.1:18093"));
1269
  }
1270
 
 
 
 
 
 
 
 
 
 
 
1271
  #[test]
1272
  fn only_loopback_hosts() {
1273
  for h in ["127.0.0.1", "127.0.0.1:18093", "localhost:18093", "LOCALHOST", "[::1]:18093", "[::1]"] {
 
43
  pub registry: Option<PathBuf>,
44
  /// `--native`: how long `/api/*` keeps a model resident when the request does not say (Ollama's default, 5m)
45
  pub keep_alive: Option<String>,
46
+ /// one web page origin (`https://host[:port]`) whose scripts may call this gateway from a browser: CORS and
47
+ /// Chrome's private-network preflight are answered for it alone; the loopback `Host` rule still holds
48
+ pub allow_origin: Option<String>,
49
+ }
50
+
51
+ /// Whether `o` is an origin: `http(s)://host[:port]`, no path, query or credentials.
52
+ pub fn is_origin(o: &str) -> bool {
53
+ let rest = o.strip_prefix("https://").or_else(|| o.strip_prefix("http://"));
54
+ rest.is_some_and(|r| !r.is_empty() && r.chars().all(|c| c.is_ascii_alphanumeric() || matches!(c, '.' | '-' | ':' | '[' | ']')))
55
+ }
56
+
57
+ thread_local! {
58
+ /// The CORS headers for the request this connection thread is answering (empty unless its `Origin` is allowed).
59
+ static CORS: std::cell::RefCell<String> = const { std::cell::RefCell::new(String::new()) };
60
+ }
61
+
62
+ /// The CORS header lines to add to this connection's response (`\r\n`-terminated; empty when none apply).
63
+ pub(crate) fn cors_headers() -> String {
64
+ CORS.with(|c| c.borrow().clone())
65
  }
66
 
67
  struct State {
 
76
  ident: FileIdent,
77
  /// where `/slots/{id}?action=save|restore` keeps slot files (`--slot-dir`; native mode)
78
  slot_dir: Option<PathBuf>,
79
+ /// `--allow-origin`: the one web origin answered with CORS headers
80
+ allow_origin: Option<String>,
81
  }
82
 
83
  /// The file identity checked before each P0 answer: device, inode, size and mtime, not the bytes (re-hashing per
 
131
 
132
  /// Verify the model, launch or check the upstream (or start `--native`), then serve until killed.
133
  pub fn run(cfg: Config) -> Result<(), String> {
134
+ if let Some(o) = cfg.allow_origin.as_deref().filter(|o| !is_origin(o)) {
135
+ return Err(format!("--allow-origin {o}: an origin is http(s)://host[:port], with no path"));
136
+ }
137
  let before = ident(&cfg.model).map_err(|e| format!("{}: {e}", cfg.model.display()))?;
138
  let verified = crate::verify(&cfg.model, &cfg.fork_json, cfg.engine)?;
139
  let model = cfg.model.canonicalize().map_err(|e| format!("{}: {e}", cfg.model.display()))?;
 
177
  }
178
  let hashed_at = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
179
  let engine = "llama.cpp b11192 llama-server (loopback), behind bankml P0".to_string();
180
+ let st = Arc::new(State { native: None, keep_alive: crate::native::KeepAlive::Forever, verified, model, upstream, engine, hashed_at, ident: id, slot_dir: None, allow_origin: cfg.allow_origin.clone() });
181
  let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
182
  eprintln!("bankml serve {}: {} verified (sha256 {}), upstream {} serves it; listening on http://{}",
183
  crate::VERSION, st.model.display(), st.verified.model_sha256, st.upstream, cfg.listen);
 
213
  if let Some(d) = &cfg.slot_dir {
214
  std::fs::create_dir_all(d).map_err(|e| format!("--slot-dir {}: {e}", d.display()))?;
215
  }
216
+ let st = Arc::new(State { slot_dir: cfg.slot_dir.clone(), allow_origin: cfg.allow_origin.clone(), native: Some(rs), keep_alive: ka, verified: Verified { model_sha256: sha.clone(), guard: "play", engine: cfg.engine.as_str(), arch: None, name: None, types: Vec::new() },
217
  model, upstream: upstream.clone(), engine, hashed_at, ident: id });
218
  let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
219
  let lu = TcpListener::bind(&upstream).map_err(|e| format!("cannot listen on the engine address {upstream}: {e} (is llama-server running there?)"))?;
 
402
  _ => "Error",
403
  };
404
  let code = if (100..600).contains(&code) { code } else { 502 }; // a garbled upstream head is a bad gateway
405
+ write!(c, "HTTP/1.1 {code} {reason}\r\nContent-Type: {ctype}\r\nContent-Length: {}\r\n{}Connection: close\r\n\r\n", body.len(), cors_headers())?;
406
  c.write_all(body)
407
  }
408
 
 
429
  if !loopback_host(header(&h, "host").unwrap_or("")) {
430
  return refuse(403, b"bankml serve answers loopback clients only (Host must be 127.0.0.1, localhost or [::1])", &mut r);
431
  }
432
+ // --allow-origin: that one web page may read the answers from a browser; every other origin gets no CORS headers
433
+ let allowed = header(&h, "origin").filter(|o| st.allow_origin.as_deref() == Some(*o)).map(str::to_string);
434
+ CORS.with(|c| *c.borrow_mut() = allowed.as_deref().map(|o| format!("Access-Control-Allow-Origin: {o}\r\nVary: Origin\r\n")).unwrap_or_default());
435
+ if method == "OPTIONS" {
436
+ if allowed.is_none() {
437
+ return refuse(403, b"bankml serve: this origin is not allowed (start serve with --allow-origin ORIGIN)", &mut r);
438
+ }
439
+ // the preflight; Chrome asks before a public page may reach a loopback address (Private Network Access)
440
+ let pna = header(&h, "access-control-request-private-network").is_some_and(|v| v.eq_ignore_ascii_case("true"));
441
+ return write!(c, "HTTP/1.1 204 No Content\r\n{}Access-Control-Allow-Methods: GET, POST\r\nAccess-Control-Allow-Headers: Content-Type\r\n{}Access-Control-Max-Age: 600\r\nContent-Length: 0\r\nConnection: close\r\n\r\n",
442
+ cors_headers(), if pna { "Access-Control-Allow-Private-Network: true\r\n" } else { "" });
443
+ }
444
  if header(&h, "transfer-encoding").is_some() {
445
  return refuse(400, b"chunked requests are not accepted; send Content-Length", &mut r);
446
  }
 
619
  let model_id = l.model.file_name().map(|n| n.to_string_lossy().into_owned()).unwrap_or_default();
620
  let created = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
621
  let r = if stream {
622
+ write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\n{}Connection: close\r\n\r\n", cors_headers())?;
623
  // with logprobs, llama-server's shape: the first token's partial opens with the role delta, and each token's
624
  // entry rides on the last delta its partial sends (no delta, no entry)
625
  let logprobs = nc.params.n_probs > 0;
 
1014
  let merged = format!("{}, \"bankml_receipt\": {}}}", &s[..end], t.receipt(&st.engine, &st.verified));
1015
  return respond(c, 200, "application/json", merged.as_bytes());
1016
  }
1017
+ write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\n{}Connection: close\r\n\r\n", cors_headers())?;
1018
  let mut pending = Vec::new();
1019
  while let Some(ch) = b.next(&mut r)? {
1020
  pending.extend(ch);
 
1304
  assert_eq!(header(&h, "host"), Some("127.0.0.1:18093"));
1305
  }
1306
 
1307
+ #[test]
1308
+ fn origins_are_scheme_host_port_only() {
1309
+ for o in ["https://pythai-bankml.static.hf.space", "http://127.0.0.1:8000", "https://[::1]:7860"] {
1310
+ assert!(is_origin(o), "{o}");
1311
+ }
1312
+ for o in ["https://a.example/path", "pythai-bankml.static.hf.space", "https://", "https://u@h", "https://h?x=1", "*"] {
1313
+ assert!(!is_origin(o), "{o}");
1314
+ }
1315
+ }
1316
+
1317
  #[test]
1318
  fn only_loopback_hosts() {
1319
  for h in ["127.0.0.1", "127.0.0.1:18093", "localhost:18093", "LOCALHOST", "[::1]:18093", "[::1]"] {
bankml-chat.js ADDED
@@ -0,0 +1,204 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ // SPDX-License-Identifier: MIT OR Apache-2.0
2
+ // Ask bankML, from this static page: two ways, never confused.
3
+ // · Your own bankML — this browser talks to `bankml serve` on the visitor's machine (started with
4
+ // --allow-origin for this page): bankML's verified arithmetic, the visitor's CPU and RAM, a receipt on every
5
+ // answer whose sha256 this page checks. Free.
6
+ // · A Hugging Face provider — sign in with Hugging Face; a provider-hosted model answers under bankML's persona,
7
+ // billed to the visitor's inference quota. NOT bankML's arithmetic, no receipt; said on every answer.
8
+ // Nothing an engine or a provider returns becomes markup: text only.
9
+ import { oauthLoginUrl, oauthHandleRedirectIfPresent } from "https://cdn.jsdelivr.net/npm/@huggingface/hub@2.17.1/+esm";
10
+ import { InferenceClient } from "https://cdn.jsdelivr.net/npm/@huggingface/inference@4.13.28/+esm";
11
+
12
+ const $ = (id) => document.getElementById(id);
13
+ const SESSION = "bankml.oauth";
14
+ const DEFAULT_ENDPOINT = "http://127.0.0.1:18093";
15
+ const DEFAULT_MODEL = "Qwen/Qwen3-8B";
16
+ const history = [];
17
+ let persona = null;
18
+ let oauth = null;
19
+
20
+ const mode = () => document.querySelector('input[name="mode"]:checked').value;
21
+
22
+ // ── the persona: the same file the bankML console speaks from ─────────────────────────────────────────────────
23
+ async function loadPersona() {
24
+ try {
25
+ const r = await fetch("sAGI/personas/bankml.persona", { cache: "no-store" });
26
+ persona = await r.json();
27
+ } catch (e) {
28
+ persona = null;
29
+ $("askstatus").textContent = "the persona could not be read: " + e;
30
+ }
31
+ }
32
+
33
+ // ── the SELF block, as the console builds it: one sentence per measurement, "not measured" when it was not ──────
34
+ const get = async (ep, path) => {
35
+ try {
36
+ const r = await fetch(ep + path, { cache: "no-store" });
37
+ return r.ok ? await r.json() : null;
38
+ } catch { return null; }
39
+ };
40
+ async function selfText(ep) {
41
+ const [b, u, m] = [await get(ep, "/bankml") || {}, await get(ep, "/bankml/usage") || {}, await get(ep, "/bankml/metrics") || {}];
42
+ const last = (m.records || [{}]).slice(-1)[0] || {};
43
+ const v = (x, unit = "", scale = 1, d = 1) => (x === null || x === undefined ? "not measured" : (x * scale).toFixed(d) + unit);
44
+ const ver = b.verified || {};
45
+ return [
46
+ `- the model I am running: ${ver.name || b.resident || "none"} (sha256 ${String(ver.model_sha256 || "").slice(0, 16)}…), bankML ${ver.bankml}`,
47
+ `- tokens I have read in total (prompts): ${v(m.prompt_tokens, "", 1, 0)}`,
48
+ `- tokens I have written in total (answers): ${v(m.completion_tokens, "", 1, 0)}`,
49
+ `- time to first token of my last answer: ${v(last.ttft_ms, " milliseconds", 1, 0)}`,
50
+ `- prompt reading speed of my last answer: ${v(last.prompt_tps, " tokens per second")}`,
51
+ `- generation speed of my last answer: ${v(last.eval_tps, " tokens per second")}`,
52
+ `- CPU use right now: ${v(u.cpu_percent, " percent of one core")}`,
53
+ `- memory I hold (RSS): ${v(u.rss_bytes, " GB", 1e-9, 2)}; memory still available on this machine: ${v(u.mem_available_bytes, " GB", 1e-9, 2)}`,
54
+ `- power the CPU package draws: ${v(u.package_watts, " watts")}; energy per token I write: ${v(m.joules_per_token, " joules", 1, 3)}`,
55
+ ].join("\n");
56
+ }
57
+
58
+ // ── the conversation ──────────────────────────────────────────────────────────────────────────────────────────
59
+ function bubble(role, text) {
60
+ const d = document.createElement("div");
61
+ d.className = "msg " + role;
62
+ d.textContent = text;
63
+ $("chat").append(d);
64
+ d.scrollIntoView({ block: "nearest" });
65
+ return d;
66
+ }
67
+ function note(el, text, cls) {
68
+ const p = document.createElement("div");
69
+ p.className = "meta " + (cls || "");
70
+ p.textContent = text;
71
+ el.after(p);
72
+ return p;
73
+ }
74
+ async function sha256hex(text) {
75
+ const h = await crypto.subtle.digest("SHA-256", new TextEncoder().encode(text));
76
+ return [...new Uint8Array(h)].map((b) => b.toString(16).padStart(2, "0")).join("");
77
+ }
78
+
79
+ // ── your own bankML ───────────────────────────────────────────────────────────────────────────────────────────
80
+ async function connect() {
81
+ const ep = $("endpoint").value.trim().replace(/\/+$/, "") || DEFAULT_ENDPOINT;
82
+ $("localstatus").textContent = "connecting…";
83
+ const b = await get(ep, "/bankml");
84
+ if (!b) {
85
+ $("localstatus").textContent = "not reachable — is bankml serve running with --allow-origin " + location.origin + " ? (see below)";
86
+ return false;
87
+ }
88
+ const v = b.verified || {};
89
+ $("localstatus").textContent = v.guard === "play"
90
+ ? `✓ connected: ${v.name || b.resident || "model"} · bankML ${v.bankml} · sha256 ${String(v.model_sha256 || "").slice(0, 12)}…`
91
+ : "connected, but no verified model is loaded yet";
92
+ return v.guard === "play";
93
+ }
94
+ async function askLocal(message) {
95
+ const ep = $("endpoint").value.trim().replace(/\/+$/, "") || DEFAULT_ENDPOINT;
96
+ const system = persona.system_prompt + "\n\nSELF (measured by bankML just now):\n" + await selfText(ep);
97
+ const out = bubble("assistant", "…");
98
+ let text = "", receipt = null;
99
+ const r = await fetch(ep + "/v1/chat/completions", {
100
+ method: "POST", headers: { "Content-Type": "application/json" },
101
+ body: JSON.stringify({ messages: [{ role: "system", content: system }, ...history.slice(-12), { role: "user", content: message }],
102
+ stream: true, max_tokens: 384 }),
103
+ });
104
+ if (!r.ok) throw new Error(`HTTP ${r.status}: ${(await r.text()).slice(0, 300)}`);
105
+ const reader = r.body.getReader(), dec = new TextDecoder();
106
+ let buf = "";
107
+ for (;;) {
108
+ const { done, value } = await reader.read();
109
+ if (done) break;
110
+ buf += dec.decode(value, { stream: true });
111
+ let i;
112
+ while ((i = buf.indexOf("\n")) >= 0) {
113
+ const line = buf.slice(0, i).trim();
114
+ buf = buf.slice(i + 1);
115
+ if (!line.startsWith("data:") || line === "data: [DONE]") continue;
116
+ const d = JSON.parse(line.slice(5));
117
+ if (d.bankml_receipt) { receipt = d.bankml_receipt; continue; }
118
+ const piece = d.choices?.[0]?.delta?.content;
119
+ if (piece) { text += piece; out.textContent = text; }
120
+ }
121
+ }
122
+ const ok = receipt && (await sha256hex(text)) === receipt.response_sha256;
123
+ note(out, receipt
124
+ ? `${ok ? "✓" : "✗"} receipt — bankML ${receipt.bankml} · model sha256 ${String(receipt.model_sha256 || "").slice(0, 12)}… · answer sha256 ${ok ? "matches the text received" : "does NOT match the text received"}`
125
+ : "no receipt came with this answer", ok ? "ok" : "bad");
126
+ history.push({ role: "user", content: message }, { role: "assistant", content: text });
127
+ }
128
+
129
+ // ── a Hugging Face provider (not bankML) ──────────────────────────────────────────────────────────────────────
130
+ function readSession() { try { return JSON.parse(sessionStorage.getItem(SESSION) || "null"); } catch { return null; } }
131
+ function writeSession(v) { try { v ? sessionStorage.setItem(SESSION, JSON.stringify(v)) : sessionStorage.removeItem(SESSION); } catch {} }
132
+ const signedIn = () => !!(oauth && oauth.accessToken && new Date(oauth.accessTokenExpiresAt) > new Date());
133
+ function renderAuth() {
134
+ $("signin").hidden = signedIn();
135
+ $("signout").hidden = !signedIn();
136
+ $("who").textContent = signedIn()
137
+ ? `signed in as ${oauth.userInfo?.preferred_username || oauth.userInfo?.name || "you"} — answers spend your inference quota`
138
+ : "not signed in";
139
+ }
140
+ async function askProvider(message) {
141
+ const model = $("model").value.trim() || DEFAULT_MODEL;
142
+ if (!/^[\w.-]+\/[\w.-]+$/.test(model)) throw new Error("the model must be a Hugging Face repository id, owner/name");
143
+ // the persona, told the truth about where it is running
144
+ const system = persona.system_prompt + `\n\nIMPORTANT: this answer is NOT produced by bankML. It is produced by ${model} through a Hugging Face inference provider: there is no bankML arithmetic, no receipt and no SELF block. Say so if you are asked about your speed, your use, your receipt or your verification.`;
145
+ const out = bubble("assistant", "…");
146
+ let text = "";
147
+ const client = new InferenceClient(oauth.accessToken);
148
+ const stream = client.chatCompletionStream({ provider: "auto", model, max_tokens: 384,
149
+ messages: [{ role: "system", content: system }, ...history.slice(-12), { role: "user", content: message }] });
150
+ for await (const chunk of stream) {
151
+ const piece = chunk?.choices?.[0]?.delta?.content;
152
+ if (piece) { text += piece; out.textContent = text; }
153
+ }
154
+ note(out, `not bankML — ${model} via a Hugging Face provider · no receipt · your quota`, "warn");
155
+ history.push({ role: "user", content: message }, { role: "assistant", content: text });
156
+ }
157
+
158
+ // ── wiring ─────────────────────────────────────────────────────────────────────────────────────────────────────
159
+ function renderMode() {
160
+ const m = mode();
161
+ $("localrow").hidden = m !== "local";
162
+ $("hfrow").hidden = m !== "hf";
163
+ $("modenote").textContent = m === "local"
164
+ ? "bankML's own arithmetic on your machine: verified model, receipt checked here, nothing sent anywhere else."
165
+ : "Not bankML: a provider-hosted model speaks with bankML's persona, without its arithmetic or a receipt.";
166
+ }
167
+ document.querySelectorAll('input[name="mode"]').forEach((r) => r.addEventListener("change", renderMode));
168
+ $("connect").addEventListener("click", connect);
169
+ $("signin").addEventListener("click", async () => {
170
+ if (!window.huggingface?.variables?.OAUTH_CLIENT_ID) { $("who").textContent = "sign-in works only on the Hugging Face Space itself"; return; }
171
+ window.location.href = await oauthLoginUrl({ scopes: window.huggingface.variables.OAUTH_SCOPES });
172
+ });
173
+ $("signout").addEventListener("click", () => { writeSession(null); oauth = null; renderAuth(); });
174
+ $("send").addEventListener("click", async () => {
175
+ const message = $("message").value.trim();
176
+ if (!message || !persona) return;
177
+ if (mode() === "hf" && !signedIn()) { $("askstatus").textContent = "sign in with Hugging Face first"; return; }
178
+ $("message").value = "";
179
+ $("send").disabled = true;
180
+ $("askstatus").textContent = "";
181
+ bubble("user", message);
182
+ try {
183
+ await (mode() === "local" ? askLocal(message) : askProvider(message));
184
+ } catch (e) {
185
+ const local = mode() === "local";
186
+ $("askstatus").textContent = (local ? "your bankML did not answer: " : "the provider did not answer: ") + String(e.message || e).slice(0, 300)
187
+ + (local ? " — is bankml serve running with --allow-origin " + location.origin + " ?" : "");
188
+ } finally {
189
+ $("send").disabled = false;
190
+ }
191
+ });
192
+ $("message").addEventListener("keydown", (e) => { if (e.key === "Enter" && (e.ctrlKey || e.metaKey)) $("send").click(); });
193
+ $("origin").textContent = location.origin;
194
+
195
+ (async () => {
196
+ $("endpoint").value = DEFAULT_ENDPOINT;
197
+ $("model").value = DEFAULT_MODEL;
198
+ renderMode();
199
+ await loadPersona();
200
+ if (persona) $("mantra").textContent = "“" + persona.mantra + "”";
201
+ try { const res = await oauthHandleRedirectIfPresent(); if (res) writeSession(res); } catch (e) { $("who").textContent = "sign-in failed: " + String(e).slice(0, 120); }
202
+ oauth = readSession();
203
+ renderAuth();
204
+ })();
docs/BUILD_HISTORY.md CHANGED
@@ -107,7 +107,7 @@ then memory traffic per token, with strategies from llama.cpp, vLLM, Ollama and
107
  dependency justified in `Cargo.toml` comments. A single static binary.
108
  3. **Verified response** — every answer carries its evidence, and a model that cannot be
109
  verified does not answer:
110
- - before load: the GGUF guard (port of minaiml `gguf_guard.py`) → `play | refuse | need_more`;
111
  refuse is shown with its reason, never a silent fallback (Bonsai 2 Q2_0 loads on mainline
112
  and answers in gibberish — the worst failure);
113
  - the file's sha256 must match a pinned record (`FORK.json` of the PYTHAI fork);
 
107
  dependency justified in `Cargo.toml` comments. A single static binary.
108
  3. **Verified response** — every answer carries its evidence, and a model that cannot be
109
  verified does not answer:
110
+ - before load: the GGUF guard (port of [minaiml](https://github.com/minaiml) `gguf_guard.py`) → `play | refuse | need_more`;
111
  refuse is shown with its reason, never a silent fallback (Bonsai 2 Q2_0 loads on mainline
112
  and answers in gibberish — the worst failure);
113
  - the file's sha256 must match a pinned record (`FORK.json` of the PYTHAI fork);
docs/index.html CHANGED
@@ -113,6 +113,9 @@ article p[align="center"] { text-align: center; }
113
  .picker select { width: 100%; font: 15px var(--f-body); padding: 8px; border-radius: 6px; border: 1px solid var(--rule); background: var(--panel); color: var(--ink); }
114
  article h1 { font-size: 1.65rem; }
115
  }
 
 
 
116
  @media (prefers-reduced-motion: no-preference) { html { scroll-behavior: smooth; } }
117
  </style>
118
 
@@ -120,6 +123,7 @@ article p[align="center"] { text-align: center; }
120
  <div class="masthead-inner">
121
  <p class="brand"><span class="trits" aria-hidden="true"><i></i><i></i><i></i></span>bankml<span class="ver" id="ver">0.2.7</span></p>
122
  <span class="tag">Verified low-bit inference for the CPU you already have — the documentation, as of <span id="asof">0.2.7</span>.</span>
 
123
  <div class="proof">gate record 0.3.6: the penalties — repeat, frequency and presence, token-identical to llama-server on <b>56 / 56</b> answers on each of three models, live <b>85 / 85</b> · gate record 0.3.5: JSON schemas — llama-server’s grammar on each template (<b>173 / 173</b> schemas × 3 templates against llama.cpp’s own code) and its answers token-identical on all five native models (<b>28 / 28</b>, <b>11 / 11</b>, <b>56 / 56</b> × 3), the content rule matched on <b>30,063 / 30,063</b> texts per template · <code>bankml create</code>: mindX’s persona layer, made from mindXtrain’s merged output, token-identical end to end (<b>27 / 27</b>, two ways) · gate record 0.3.4: mindX’s own model natively — <code>mindx-gen39</code> and SmolLM2-135M-Instruct (the Llama graph in F16) and Bonsai-1.7B (tied embeddings) token-identical to llama-server b11192: whole model <b>800 / 800</b> and <b>840 / 840</b> rows bit-exact, F16 products <b>552,268 / 552,268</b> elements bit-exact against ggml, seeded sampling <b>40 / 40</b> on each, conversations <b>9 / 9</b>, JSON mode <b>23 / 23</b> · JSON mode — answers under llama-server’s own <code>json_object</code> grammar token-identical, greedy and seeded (<b>23 / 23</b> answers on the 1-bit model, <b>13 / 13</b> on the ternary), and llama.cpp’s grammar sampler matched on <b>1,645 / 1,645</b> whole-vocabulary masks · a C API, <code>libbankml</code>, identical to <code>serve --native</code> · Savante answered by bankML’s own forward pass — whole conversations identical to llama-server on <b>9 / 9</b> turns (text, token counts, prompt-cache reuse) · own forward pass <b>1,064 / 1,064</b> rows bit-exact for the 1-bit and the ternary model · seeded sampling identical on <b>40 / 40</b> continuations · <b>8,188,239,872</b> ternary weights and <b>762 / 762</b> dot products bit-exact against llama.cpp b11192’s compiled library · one ternary token <b>0.222 s vs 2.139 s</b></div>
124
  </div>
125
  </header>
 
113
  .picker select { width: 100%; font: 15px var(--f-body); padding: 8px; border-radius: 6px; border: 1px solid var(--rule); background: var(--panel); color: var(--ink); }
114
  article h1 { font-size: 1.65rem; }
115
  }
116
+ .masthead .links { margin: 6px 0 0; font: 500 12.5px/1.4 var(--f-mono); }
117
+ .masthead .links a { color: inherit; text-decoration: underline; text-underline-offset: 2px; opacity: .85; }
118
+ .masthead .links a:hover { opacity: 1; }
119
  @media (prefers-reduced-motion: no-preference) { html { scroll-behavior: smooth; } }
120
  </style>
121
 
 
123
  <div class="masthead-inner">
124
  <p class="brand"><span class="trits" aria-hidden="true"><i></i><i></i><i></i></span>bankml<span class="ver" id="ver">0.2.7</span></p>
125
  <span class="tag">Verified low-bit inference for the CPU you already have — the documentation, as of <span id="asof">0.2.7</span>.</span>
126
+ <p class="links"><a href="https://github.com/cryptoAGI/bankml">GitHub</a> · <a href="https://huggingface.co/spaces/PYTHAI/bankml">Hugging Face</a> · <a href="https://deltaverse.pythai.net/bankml">why bankML</a> · <a href="#thesis">the thesis</a> · <a href="https://github.com/minaiml">minaiml</a> (<a href="https://huggingface.co/spaces/PYTHAI/minaiml">Space</a>)</p>
127
  <div class="proof">gate record 0.3.6: the penalties — repeat, frequency and presence, token-identical to llama-server on <b>56 / 56</b> answers on each of three models, live <b>85 / 85</b> · gate record 0.3.5: JSON schemas — llama-server’s grammar on each template (<b>173 / 173</b> schemas × 3 templates against llama.cpp’s own code) and its answers token-identical on all five native models (<b>28 / 28</b>, <b>11 / 11</b>, <b>56 / 56</b> × 3), the content rule matched on <b>30,063 / 30,063</b> texts per template · <code>bankml create</code>: mindX’s persona layer, made from mindXtrain’s merged output, token-identical end to end (<b>27 / 27</b>, two ways) · gate record 0.3.4: mindX’s own model natively — <code>mindx-gen39</code> and SmolLM2-135M-Instruct (the Llama graph in F16) and Bonsai-1.7B (tied embeddings) token-identical to llama-server b11192: whole model <b>800 / 800</b> and <b>840 / 840</b> rows bit-exact, F16 products <b>552,268 / 552,268</b> elements bit-exact against ggml, seeded sampling <b>40 / 40</b> on each, conversations <b>9 / 9</b>, JSON mode <b>23 / 23</b> · JSON mode — answers under llama-server’s own <code>json_object</code> grammar token-identical, greedy and seeded (<b>23 / 23</b> answers on the 1-bit model, <b>13 / 13</b> on the ternary), and llama.cpp’s grammar sampler matched on <b>1,645 / 1,645</b> whole-vocabulary masks · a C API, <code>libbankml</code>, identical to <code>serve --native</code> · Savante answered by bankML’s own forward pass — whole conversations identical to llama-server on <b>9 / 9</b> turns (text, token counts, prompt-cache reuse) · own forward pass <b>1,064 / 1,064</b> rows bit-exact for the 1-bit and the ternary model · seeded sampling identical on <b>40 / 40</b> continuations · <b>8,188,239,872</b> ternary weights and <b>762 / 762</b> dot products bit-exact against llama.cpp b11192’s compiled library · one ternary token <b>0.222 s vs 2.139 s</b></div>
128
  </div>
129
  </header>
docs/install.md CHANGED
@@ -353,6 +353,7 @@ oracles all do this.
353
  | `--ctx N` | `4096` | both | the context: `-c N` for the spawned llama-server, `n_ctx` for the native engine. A prompt of `N` tokens or more is refused (0.3.8: with llama-server's `exceed_context_size_error` 400 body on `/v1`); an Ollama `num_ctx` above it is refused |
354
  | `--spec-ngram` | off | llama.cpp | `--spec-type ngram-simple` in the spawned engine: n-gram speculative decoding, exact at temperature 0, opt-in because it measured within noise |
355
  | `--slot-dir DIR` | — | both | llama.cpp mode: `--slot-save-path DIR` for the spawned engine. Native mode (0.3.8): `POST /slots/0?action=save\|restore\|erase` with `{"filename"}`, as llama-server's slot API; the directory is created at start. Either way, a restart restores the system prompt instead of computing it again. Without it, the slot actions answer 501 |
 
356
  | `--native` | off | — | native mode |
357
  | `--registry [DIR]` | off; bare: `$BANKML_FORKS`, else `~/.local/share/bankml/forks` | native | serve every model pinned in `DIR` by name, one resident at a time, each verified again when it loads; answer `/api/create`, `/api/copy`, `/api/delete` for derived models. A following argument that starts with `--` is not taken as `DIR` |
358
  | `--keep-alive DUR` | `5m` | native | how long an `/api/*` request that names no `keep_alive` keeps the model resident: `"5m"`, `"1h30m"`, seconds, `0` (unload after the answer), negative (for good). OpenAI and llama-server endpoints keep the model resident |
 
353
  | `--ctx N` | `4096` | both | the context: `-c N` for the spawned llama-server, `n_ctx` for the native engine. A prompt of `N` tokens or more is refused (0.3.8: with llama-server's `exceed_context_size_error` 400 body on `/v1`); an Ollama `num_ctx` above it is refused |
354
  | `--spec-ngram` | off | llama.cpp | `--spec-type ngram-simple` in the spawned engine: n-gram speculative decoding, exact at temperature 0, opt-in because it measured within noise |
355
  | `--slot-dir DIR` | — | both | llama.cpp mode: `--slot-save-path DIR` for the spawned engine. Native mode (0.3.8): `POST /slots/0?action=save\|restore\|erase` with `{"filename"}`, as llama-server's slot API; the directory is created at start. Either way, a restart restores the system prompt instead of computing it again. Without it, the slot actions answer 501 |
356
+ | `--allow-origin ORIGIN` | — | both | 0.3.9: one web page (`https://host[:port]`) whose scripts may call this gateway from a browser — the bankML Space's page, for example (`https://pythai-bankml.static.hf.space`). bankML answers that origin's CORS preflight (with Chrome's `Access-Control-Allow-Private-Network`) and adds `Access-Control-Allow-Origin` to its answers, streamed ones included; any other origin gets no CORS headers and a 403 preflight. The loopback `Host` and JSON-POST rules are unchanged |
357
  | `--native` | off | — | native mode |
358
  | `--registry [DIR]` | off; bare: `$BANKML_FORKS`, else `~/.local/share/bankml/forks` | native | serve every model pinned in `DIR` by name, one resident at a time, each verified again when it loads; answer `/api/create`, `/api/copy`, `/api/delete` for derived models. A following argument that starts with `--` is not taken as `DIR` |
359
  | `--keep-alive DUR` | `5m` | native | how long an `/api/*` request that names no `keep_alive` keeps the model resident: `"5m"`, `"1h30m"`, seconds, `0` (unload after the answer), negative (for good). OpenAI and llama-server endpoints keep the model resident |
docs/modules/gguf.md CHANGED
@@ -3,7 +3,7 @@
3
  ## Summary
4
 
5
  `gguf.rs` reads the header of a GGUF v3 file and decides whether bankml may play it. The decision is the
6
- **guard**: `play`, `refuse` with reasons, or `need_more` header bytes. It is a port of minaiml's Python
7
  `gguf_guard.py` (same verdicts, same reasons, same JSON keys). The guard never reads tensor data.
8
 
9
  The module also holds what the rest of bankml needs to reach tensor data safely: type names, block layouts,
 
3
  ## Summary
4
 
5
  `gguf.rs` reads the header of a GGUF v3 file and decides whether bankml may play it. The decision is the
6
+ **guard**: `play`, `refuse` with reasons, or `need_more` header bytes. It is a port of [minaiml](https://github.com/minaiml)'s Python
7
  `gguf_guard.py` (same verdicts, same reasons, same JSON keys). The guard never reads tensor data.
8
 
9
  The module also holds what the rest of bankml needs to reach tensor data safely: type names, block layouts,
docs/thesis.md CHANGED
@@ -132,7 +132,76 @@ bankML's discipline is to establish exactly that, against an independent referen
132
  integrity, not proof: they bind the answer's text to a pinned file and to the request, between a client and its own
133
  gateway ([research.md §3](research.md#3-verifiable-and-attested-inference)).
134
 
135
- ### II.5 What bankML inherits, extends and breaks
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
136
 
137
  It **inherits** ggml's formats, its kernels' semantics and llama-server's protocol. It **extends** them with an x86
138
  ternary kernel the reference lacks, a gate in front of every answer, and receipts. It **breaks** with the convention
@@ -164,7 +233,8 @@ already shown identical.
164
  result and not shipped ([`testing/experiments/`](../testing/experiments/)).
165
  2. **The compiled reference is the specification, not its source.** Where the shipped library and its C source
166
  disagree, the library wins (§II.3).
167
- 3. **A model that cannot be verified does not answer.** The guard ([`bankML/gguf.rs`](../bankML/gguf.rs)) refuses,
 
168
  from the header alone and with a reason, the three low-bit traps, including a file mainline loads and answers in
169
  fluent nonsense. The pin ([`bankML/sha256.rs`](../bankML/sha256.rs)) refuses a file whose hash differs from its
170
  provenance record.
@@ -313,6 +383,8 @@ Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Cameron, P., Jaggi, M., Alis
313
 
314
  Cankaya, E. (2026). "Bit-Exact AI Inference Verification Without Performance Tradeoffs." [arXiv:2606.00279](https://arxiv.org/abs/2606.00279).
315
 
 
 
316
  Codephreak, Professor and Magnusson, G. L. (2026). Design directives for bankML and the mindX runtime, recorded in the
317
  project (project record, 2026-07-04 to 2026-09-28); quoted in [TECHNICAL.md](TECHNICAL.md#thesis--professor-codephreak-and-gregory-l-magnusson).
318
 
@@ -335,6 +407,9 @@ Surveys* 23(1): 5–48. [doi:10.1145/103162.103163](https://doi.org/10.1145/1031
335
  Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R. and Bengio, Y. (2016). "Binarized Neural Networks." *Advances in
336
  Neural Information Processing Systems 29*. [Proceedings](https://papers.nips.cc/paper_files/paper/2016/hash/d8330f857a17c53d217014ee776bfd50-Abstract.html); preprint [arXiv:1602.02830](https://arxiv.org/abs/1602.02830).
337
 
 
 
 
338
  Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H. and Stoica, I. (2023).
339
  "Efficient Memory Management for Large Language Model Serving with PagedAttention." *Proceedings of the 29th Symposium
340
  on Operating Systems Principles (SOSP 2023)*. [arXiv:2309.06180](https://arxiv.org/abs/2309.06180).
@@ -344,9 +419,21 @@ Software* 39(2). [doi:10.1109/MS.2021.3073045](https://doi.org/10.1109/MS.2021.3
344
 
345
  Li, F., Zhang, B. and Liu, B. (2016). "Ternary Weight Networks." [arXiv:1605.04711](https://arxiv.org/abs/1605.04711).
346
 
 
 
 
 
 
 
 
 
347
  Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J. and Wei, F. (2024). "The Era of
348
  1-bit LLMs: All Large Language Models are in 1.58 Bits." [arXiv:2402.17764](https://arxiv.org/abs/2402.17764).
349
 
 
 
 
 
350
  Qwen Team (2025). "Qwen3 Technical Report." [arXiv:2505.09388](https://arxiv.org/abs/2505.09388).
351
 
352
  Rastegari, M., Ordonez, V., Redmon, J. and Farhadi, A. (2016). "XNOR-Net: ImageNet Classification Using Binary
@@ -360,6 +447,14 @@ Thompson, K. (1984). "Reflections on Trusting Trust." *Communications of the ACM
360
  Tseng, A., Chee, J., Sun, Q., Kuleshov, V. and De Sa, C. (2024). "QuIP#: Even Better LLM Quantization with Hadamard
361
  Incoherence and Lattice Codebooks." *International Conference on Machine Learning (ICML 2024)*. [arXiv:2402.04396](https://arxiv.org/abs/2402.04396).
362
 
 
 
 
 
 
 
 
 
363
  Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y. and Wei, F. (2023). "BitNet:
364
  Scaling 1-bit Transformers for Large Language Models." [arXiv:2310.11453](https://arxiv.org/abs/2310.11453).
365
 
@@ -372,9 +467,18 @@ Wang, J., Zhou, H., Song, T. et al. (2025). "Bitnet.cpp: Efficient Edge Inferenc
372
  Wei, J. et al. (2025). "T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge." *Proceedings of
373
  EuroSys 2025*. [arXiv:2407.00088](https://arxiv.org/abs/2407.00088).
374
 
 
 
 
 
 
 
 
375
  Zhu, C., Han, S., Mao, H. and Dally, W. J. (2017). "Trained Ternary Quantization." *International Conference on
376
  Learning Representations (ICLR 2017)*. [arXiv:1612.01064](https://arxiv.org/abs/1612.01064).
377
 
 
 
378
  *Notes on the references.* Author lists for the 2024–2026 preprints follow [research.md](research.md), which records
379
  which details were re-fetched and which are as commonly cited. QuaRot and QuIP# are cited for the technique of
380
  rotating activations by a Hadamard transform before quantization; the attribution of llama.cpp's own rotation to
 
132
  integrity, not proof: they bind the answer's text to a pinned file and to the request, between a client and its own
133
  gateway ([research.md §3](research.md#3-verifiable-and-attested-inference)).
134
 
135
+ ### II.5 The contemporary field (survey of 6 October 2026)
136
+
137
+ This section maps the work bankML sits among, as found on 6 October 2026: every paper below was checked against its
138
+ arXiv abstract, and every repository's existence and last activity against GitHub on that date. Projects' own speed
139
+ figures are theirs, not reproduced here. A wider survey with the verifiable-inference field is in
140
+ [research.md](research.md).
141
+
142
+ #### II.5.1 Ternary and 1-bit models
143
+
144
+ Two programmes now train language models through the ternary constraint rather than quantizing after training. The
145
+ BitNet line (Microsoft Research) moved from the 1.58-bit recipe (Ma et al. 2024) to an openly released 2-billion-
146
+ parameter ternary model trained on 4 trillion tokens (Ma et al. 2025), then to 4-bit activations beside the 1-bit
147
+ weights (Wang, Ma and Wei 2024; 2025, the latter by a Hadamard transformation of the activations), and to distilling
148
+ full-precision models into 1.58 bits (Wu et al. 2025). Spectra (Kaushal et al. 2024) trained a suite of 54 models from
149
+ 99 M to 3.9 B parameters to compare ternary "TriLMs" with float and post-training-quantized peers at equal size, and
150
+ Spectra 1.1 (Vaidhya et al. 2025) scaled TriLMs to 1.2 T tokens with scaling laws and its own inference kernel.
151
+ Alongside these, TernaryLLM (Chen et al. 2024) and OneBit (Xu et al. 2024) push quantization-aware training of
152
+ existing models to ternary and 1-bit weights, PTQ1.61 (Zhao et al. 2025) reaches below two bits after training,
153
+ ParetoQ (Liu et al. 2025) maps scaling laws across extreme bit widths, and MatMul-free language modelling (Zhu et al.
154
+ 2024) removes matrix multiplication by ternary weights altogether.
155
+
156
+ The Bonsai models bankML serves come from PrismML and are dense Qwen3 networks trained to one bit (`Q1_0`) and to
157
+ ternary values (`Q2_0`, group 64, 2.25 bits per weight). PrismML's own format table
158
+ ([docs.prismml.com](https://docs.prismml.com/download/formats)) records that group-64 `Q2_0` is in mainline llama.cpp,
159
+ that a group-128 legacy `Q2_0` is deprecated, and that its newer Ternary Bonsai 2 formats (`PQ2_0`, `PTQ1_0`) use a
160
+ rotated weight basis that needs an activation-side Walsh–Hadamard transform at run time — the file a stock engine
161
+ loads and answers in fluent nonsense, which is why bankML's guard refuses it (§III.2), and the same family of
162
+ rotation bankML reproduced for llama.cpp's quantized cache (§III.4).
163
+
164
+ #### II.5.2 Kernels and engines for ternary and 1-bit weights
165
+
166
+ | work | what it is | relation to bankML |
167
+ |---|---|---|
168
+ | [llama.cpp](https://github.com/ggml-org/llama.cpp) upstream | `Q1_0` ([#21273](https://github.com/ggml-org/llama.cpp/pull/21273); x86 AVX2+FMA in [#21636](https://github.com/ggml-org/llama.cpp/pull/21636)); `Q2_0` ([#24448](https://github.com/ggml-org/llama.cpp/pull/24448), NEON and scalar); an x86 `Q2_0` kernel needing AVX-VNNI ([#26348](https://github.com/ggml-org/llama.cpp/pull/26348), open); the older `TQ1_0`/`TQ2_0` ([#10010](https://github.com/ggml-org/llama.cpp/pull/10010)) | the reference: bankML reproduces b11192's compiled code bit for bit and adds the plain-AVX2 `Q2_0` kernel it lacks |
169
+ | [PrismML's fork](https://github.com/PrismML-Eng/llama.cpp) | kernels for the Bonsai 2 formats, e.g. `PQ2_0` AVX2/AVX-VNNI ([#206](https://github.com/PrismML-Eng/llama.cpp/pull/206)) | different formats; refused by bankML's guard until a kernel and an oracle exist |
170
+ | [bitnet.cpp](https://github.com/microsoft/BitNet) (Wang, Zhou, Song et al. 2024; 2025) | lookup-table and I2_S kernels for BitNet b1.58 on CPU; reports 2.37–6.17× on x86 | a different weight format (BitNet's), and a different exactness claim (lossless to its own model, not to an external reference) |
171
+ | [T-MAC](https://github.com/microsoft/T-MAC) (Wei et al. 2025) | lookup-table mixed-precision GEMM on CPU and NPU | the table-lookup alternative to bankML's `maddubs` arithmetic |
172
+ | Vec-LUT (Li et al. 2025) | vector table lookup for parallel ultra-low-bit inference on edge devices | the same direction as T-MAC, parallelised |
173
+ | Spectra 1.1's TriRun (Vaidhya et al. 2025) | a GPU kernel for packed ternary weights | GPU, not CPU |
174
+
175
+ #### II.5.3 Inference engines written in Rust
176
+
177
+ | engine | what it is | state on 2026-10-06 | exactness claim |
178
+ |---|---|---|---|
179
+ | [candle](https://github.com/huggingface/candle) (Hugging Face) | a minimalist ML framework; GGUF K-quants on CPU (AVX2/NEON), CUDA, Metal, WASM | active, the base most Rust LLM projects build on | none against llama.cpp found |
180
+ | [mistral.rs](https://github.com/EricLBuehler/mistral.rs) | an LLM server on candle: many architectures, ISQ, GPTQ/AWQ/HQQ/FP8 | active | none found |
181
+ | [burn](https://github.com/tracel-ai/burn) | a general tensor and deep-learning framework | active | not a GGUF engine |
182
+ | [Crane](https://github.com/lucasjinreal/Crane), [kalosm](https://github.com/floneum/kalosm), [cake](https://github.com/evilsocket/cake) | LLM/VLM engines and libraries on candle; cake distributes inference across devices | active | none found |
183
+ | [OxiLLaMa](https://github.com/cool-japan/oxillama) | a pure-Rust GGUF engine with its own AVX2/AVX-512/NEON kernels, including `TQ1_0`/`TQ2_0` and `Q1_0_G128` | active (alpha) | top-1 logit parity within a tolerance |
184
+ | [Frink](https://github.com/antonellof/frink) | a pure-Rust GGUF engine with quantized CPU, Metal and CUDA kernels and MoE | active | quantizer bytes identical on two formats |
185
+ | [Cera](https://github.com/hyeons-lab/cera) | a Rust-native GGUF engine (AVX2/AVX-512, NEON dotprod/i8mm, optional wgpu) | crate 0.6.3, 2026-09-25 | none stated |
186
+ | [llama-gguf](https://github.com/Lexmata/llama-gguf), [lm.rs](https://github.com/samuel-vitorino/lm.rs) | small engines, correctness-first / minimal | llama-gguf 2026-04; lm.rs inactive since 2024-10 | none stated |
187
+ | [bitnet-rs](https://github.com/lilyco-42/bitnet-rs), [bitnet-toy](https://github.com/tidynest/bitnet-toy) | BitNet b1.58 in Rust (a port of bitnet.cpp; a from-scratch teaching engine) | 2026-09 | unit tests against its own C++ baseline |
188
+ | [alice-aegis](https://github.com/Aefinity-AI/alice-aegis) | a `no_std` UEFI ternary engine with frozen integer semantics and SHA-256 receipts chaining the logits | 2026-10 | bit-identical across its own ISAs, not against an external reference |
189
+ | [ratchet](https://github.com/huggingface/ratchet), [tract](https://github.com/sonos/tract), [rten](https://github.com/robertknight/rten) | browser/WebGPU and ONNX inference | active | not GGUF engines |
190
+ | [rustformers/llm](https://github.com/rustformers/llm), [llama-cpp-rs](https://github.com/utilityai/llama-cpp-rs) | the first, archived (2024); the second, bindings to llama.cpp's C++ | — | inherit llama.cpp's arithmetic by calling it |
191
+
192
+ #### II.5.4 Where bankML stands among them
193
+
194
+ Three things in this survey are bankML's alone. It is the only engine found that reproduces the reference's
195
+ *compiled* library bit for bit — every weight, every dot product, every token of whole conversations — rather than
196
+ within a tolerance or against itself; alice-aegis shares the receipt idea and the exact integer discipline but checks
197
+ against its own builds. It is the only Rust engine found with ggml's group-64 ternary `Q2_0`, and its plain-AVX2 kernel
198
+ fills the x86 gap that upstream's open VNNI kernel leaves on CPUs without VNNI. And it is one zero-dependency binary
199
+ whose every answer carries a receipt. It is behind the field in breadth: candle and mistral.rs run far more
200
+ architectures and formats and run on GPUs, OxiLLaMa and upstream llama.cpp have NEON kernels for these formats, and
201
+ the BitNet and T-MAC kernels serve a different ternary format at speeds bankML has not been compared with. The
202
+ research programmes above produce the models; bankML's contribution is to run the ones in ggml's formats exactly.
203
+
204
+ ### II.6 What bankML inherits, extends and breaks
205
 
206
  It **inherits** ggml's formats, its kernels' semantics and llama-server's protocol. It **extends** them with an x86
207
  ternary kernel the reference lacks, a gate in front of every answer, and receipts. It **breaks** with the convention
 
233
  result and not shipped ([`testing/experiments/`](../testing/experiments/)).
234
  2. **The compiled reference is the specification, not its source.** Where the shipped library and its C source
235
  disagree, the library wins (§II.3).
236
+ 3. **A model that cannot be verified does not answer.** The guard ([`bankML/gguf.rs`](../bankML/gguf.rs), a port of
237
+ the GGUF guard of [minaiml](https://github.com/minaiml), the authors' delivery layer for models on laptops and phones) refuses,
238
  from the header alone and with a reason, the three low-bit traps, including a file mainline loads and answers in
239
  fluent nonsense. The pin ([`bankML/sha256.rs`](../bankML/sha256.rs)) refuses a file whose hash differs from its
240
  provenance record.
 
383
 
384
  Cankaya, E. (2026). "Bit-Exact AI Inference Verification Without Performance Tradeoffs." [arXiv:2606.00279](https://arxiv.org/abs/2606.00279).
385
 
386
+ Chen, T., Li, Z., Xu, W. et al. (2024). "TernaryLLM: Ternarized Large Language Model." [arXiv:2406.07177](https://arxiv.org/abs/2406.07177).
387
+
388
  Codephreak, Professor and Magnusson, G. L. (2026). Design directives for bankML and the mindX runtime, recorded in the
389
  project (project record, 2026-07-04 to 2026-09-28); quoted in [TECHNICAL.md](TECHNICAL.md#thesis--professor-codephreak-and-gregory-l-magnusson).
390
 
 
407
  Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R. and Bengio, Y. (2016). "Binarized Neural Networks." *Advances in
408
  Neural Information Processing Systems 29*. [Proceedings](https://papers.nips.cc/paper_files/paper/2016/hash/d8330f857a17c53d217014ee776bfd50-Abstract.html); preprint [arXiv:1602.02830](https://arxiv.org/abs/1602.02830).
409
 
410
+ Kaushal, A., Vaidhya, T., Mondal, A. K. et al. (2024). "Spectra: Surprising Effectiveness of Pretraining Ternary
411
+ Language Models at Scale." [arXiv:2407.12327](https://arxiv.org/abs/2407.12327).
412
+
413
  Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H. and Stoica, I. (2023).
414
  "Efficient Memory Management for Large Language Model Serving with PagedAttention." *Proceedings of the 29th Symposium
415
  on Operating Systems Principles (SOSP 2023)*. [arXiv:2309.06180](https://arxiv.org/abs/2309.06180).
 
419
 
420
  Li, F., Zhang, B. and Liu, B. (2016). "Ternary Weight Networks." [arXiv:1605.04711](https://arxiv.org/abs/1605.04711).
421
 
422
+ Li, X., Yin, C., Wang, W. et al. (2025). "Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge
423
+ Devices." [arXiv:2512.06443](https://arxiv.org/abs/2512.06443).
424
+
425
+ Liu, Z., Zhao, C., Huang, H. et al. (2025). "ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization."
426
+ [arXiv:2502.02631](https://arxiv.org/abs/2502.02631).
427
+
428
+ Ma, S., Wang, H., Huang, S. et al. (2025). "BitNet b1.58 2B4T Technical Report." [arXiv:2504.12285](https://arxiv.org/abs/2504.12285).
429
+
430
  Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J. and Wei, F. (2024). "The Era of
431
  1-bit LLMs: All Large Language Models are in 1.58 Bits." [arXiv:2402.17764](https://arxiv.org/abs/2402.17764).
432
 
433
+ minaiml (software). "min ai ml — language models in miniature": delivery software for models on laptops and
434
+ phones, the origin of bankML's GGUF guard. [github.com/minaiml](https://github.com/minaiml) · [Hugging Face Space
435
+ PYTHAI/minaiml](https://huggingface.co/spaces/PYTHAI/minaiml).
436
+
437
  Qwen Team (2025). "Qwen3 Technical Report." [arXiv:2505.09388](https://arxiv.org/abs/2505.09388).
438
 
439
  Rastegari, M., Ordonez, V., Redmon, J. and Farhadi, A. (2016). "XNOR-Net: ImageNet Classification Using Binary
 
447
  Tseng, A., Chee, J., Sun, Q., Kuleshov, V. and De Sa, C. (2024). "QuIP#: Even Better LLM Quantization with Hadamard
448
  Incoherence and Lattice Codebooks." *International Conference on Machine Learning (ICML 2024)*. [arXiv:2402.04396](https://arxiv.org/abs/2402.04396).
449
 
450
+ Vaidhya, T., Kaushal, A., Jain, V. et al. (2025). "Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language
451
+ Models." [arXiv:2506.23025](https://arxiv.org/abs/2506.23025).
452
+
453
+ Wang, H., Ma, S. and Wei, F. (2024). "BitNet a4.8: 4-bit Activations for 1-bit LLMs." [arXiv:2411.04965](https://arxiv.org/abs/2411.04965).
454
+
455
+ Wang, H., Ma, S. and Wei, F. (2025). "BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs."
456
+ [arXiv:2504.18415](https://arxiv.org/abs/2504.18415).
457
+
458
  Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y. and Wei, F. (2023). "BitNet:
459
  Scaling 1-bit Transformers for Large Language Models." [arXiv:2310.11453](https://arxiv.org/abs/2310.11453).
460
 
 
467
  Wei, J. et al. (2025). "T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge." *Proceedings of
468
  EuroSys 2025*. [arXiv:2407.00088](https://arxiv.org/abs/2407.00088).
469
 
470
+ Wu, X., Huang, S., Wang, W. et al. (2025). "BitNet Distillation." [arXiv:2510.13998](https://arxiv.org/abs/2510.13998).
471
+
472
+ Xu, Y., Han, X., Yang, Z. et al. (2024). "OneBit: Towards Extremely Low-bit Large Language Models." [arXiv:2402.11295](https://arxiv.org/abs/2402.11295).
473
+
474
+ Zhao, J., Zhang, M., Wang, M. et al. (2025). "PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training
475
+ Quantization Methods for Large Language Models." [arXiv:2502.13179](https://arxiv.org/abs/2502.13179).
476
+
477
  Zhu, C., Han, S., Mao, H. and Dally, W. J. (2017). "Trained Ternary Quantization." *International Conference on
478
  Learning Representations (ICLR 2017)*. [arXiv:1612.01064](https://arxiv.org/abs/1612.01064).
479
 
480
+ Zhu, R.-J., Zhang, Y., Abreu, S. et al. (2024). "Scalable MatMul-free Language Modeling." [arXiv:2406.02528](https://arxiv.org/abs/2406.02528).
481
+
482
  *Notes on the references.* Author lists for the 2024–2026 preprints follow [research.md](research.md), which records
483
  which details were re-fetched and which are as commonly cited. QuaRot and QuIP# are cited for the technique of
484
  rotating activations by a Hadamard transform before quantization; the attribution of llama.cpp's own rotation to
docs/usage.md CHANGED
@@ -247,6 +247,9 @@ In the next release (not yet tagged); each is checked against llama-server b1119
247
  llama-server `-np 1`. Like llama-server, it keeps the states of other conversations in RAM (`BANKML_CACHE_RAM`, MiB;
248
  0 off, -1 no limit; unset, 8192 MiB but at most a quarter of the memory available at load), so a conversation that comes back after
249
  another does not recompute its whole history.
 
 
 
250
  - **A smaller conversation memory (0.3.9, in progress).** `BANKML_CACHE_TYPE=q8_0` keeps the KV cache in q8_0 instead of f16,
251
  about half the memory, so a long context fits on a small machine. It is llama.cpp's `--cache-type-k q8_0
252
  --cache-type-v q8_0` exactly, with the Hadamard rotation llama.cpp applies around a quantized cache, and gives the
 
247
  llama-server `-np 1`. Like llama-server, it keeps the states of other conversations in RAM (`BANKML_CACHE_RAM`, MiB;
248
  0 off, -1 no limit; unset, 8192 MiB but at most a quarter of the memory available at load), so a conversation that comes back after
249
  another does not recompute its whole history.
250
+ - **Your bankML from a web page (0.3.9).** `--allow-origin https://pythai-bankml.static.hf.space` lets the bankML
251
+ Space's page talk to your own `bankml serve` from your browser: your CPU, your verified model, a receipt on every
252
+ answer, nothing sent anywhere else. Only that one origin is answered with CORS headers.
253
  - **A smaller conversation memory (0.3.9, in progress).** `BANKML_CACHE_TYPE=q8_0` keeps the KV cache in q8_0 instead of f16,
254
  about half the memory, so a long context fits on a small machine. It is llama.cpp's `--cache-type-k q8_0
255
  --cache-type-v q8_0` exactly, with the Hadamard rotation llama.cpp applies around a quantized cache, and gives the
index.html CHANGED
@@ -125,6 +125,27 @@
125
  .table td:first-child{color:#fff; font-weight:500; width:34%;}
126
  .road li strong{color:rgb(var(--am));}
127
  .hope p{font-size:18px; color:#eef;}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
128
  footer{max-width:780px; margin:0 auto; padding:10px 16px 90px; text-align:center; color:var(--quiet); font-size:14px;}
129
  .social{display:flex; gap:22px; justify-content:center; align-items:center; margin:14px 0 10px;}
130
  .social a{display:inline-flex; align-items:center; gap:8px; border:none; color:var(--muted);}
@@ -147,6 +168,43 @@
147
  and 0.3.8 and 0.3.9 are being built on the way to the 0.4.0 milestone.</p>
148
  </header>
149
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
150
  <section>
151
  <h2>bankML in one paragraph</h2>
152
  <p>bankML runs language models on the computer you already have: a laptop, a small server, no graphics card
@@ -420,5 +478,6 @@
420
  mount(layer);
421
  })();
422
  </script>
 
423
  </body>
424
  </html>
 
125
  .table td:first-child{color:#fff; font-weight:500; width:34%;}
126
  .road li strong{color:rgb(var(--am));}
127
  .hope p{font-size:18px; color:#eef;}
128
+ #ask .mantra{font-size:19px;color:#fff;margin:.2em 0 .6em}
129
+ #ask .modes{display:grid;gap:6px;margin:8px 0}
130
+ #ask .modes label{color:var(--muted);cursor:pointer}
131
+ #ask .modenote{font-size:14px;color:var(--quiet);margin:.3em 0 .8em}
132
+ #ask .row{display:flex;flex-wrap:wrap;gap:8px;align-items:center;margin:8px 0}
133
+ #ask input[type=text],#ask #endpoint,#ask #model{flex:1 1 220px;min-width:0;background:rgba(0,0,0,.35);color:#fff;border:1px solid var(--rule);border-radius:8px;padding:8px 10px;font:14px ui-monospace,Menlo,monospace}
134
+ #ask textarea{flex:1 1 260px;min-width:0;background:rgba(0,0,0,.35);color:#fff;border:1px solid var(--rule);border-radius:8px;padding:9px 10px;font:15px system-ui,sans-serif;resize:vertical}
135
+ #ask button{background:rgba(var(--cy),.16);color:#fff;border:1px solid rgba(var(--cy),.5);border-radius:8px;padding:8px 14px;font:600 14px system-ui,sans-serif;cursor:pointer}
136
+ #ask button:hover{background:rgba(var(--cy),.3)}
137
+ #ask button:disabled{opacity:.5;cursor:wait}
138
+ #ask .status{font-size:13.5px;color:var(--quiet);flex:1 1 100%}
139
+ #ask details{flex:1 1 100%;font-size:14.5px;color:var(--muted)}
140
+ #ask details code{word-break:break-all}
141
+ #chat{display:grid;gap:8px;margin:12px 0 4px;max-height:420px;overflow-y:auto}
142
+ #chat .msg{white-space:pre-wrap;border-radius:10px;padding:9px 12px;line-height:1.5}
143
+ #chat .msg.user{background:rgba(var(--vi),.18);justify-self:end;max-width:85%;color:#fff}
144
+ #chat .msg.assistant{background:rgba(0,0,0,.35);border:1px solid var(--rule);color:#e8eef5}
145
+ #chat .meta{font:12.5px ui-monospace,Menlo,monospace;color:var(--quiet);margin:-4px 2px 4px}
146
+ #chat .meta.ok{color:#56d364}
147
+ #chat .meta.bad{color:#ff7b72}
148
+ #chat .meta.warn{color:rgb(var(--am))}
149
  footer{max-width:780px; margin:0 auto; padding:10px 16px 90px; text-align:center; color:var(--quiet); font-size:14px;}
150
  .social{display:flex; gap:22px; justify-content:center; align-items:center; margin:14px 0 10px;}
151
  .social a{display:inline-flex; align-items:center; gap:8px; border:none; color:var(--muted);}
 
168
  and 0.3.8 and 0.3.9 are being built on the way to the 0.4.0 milestone.</p>
169
  </header>
170
 
171
+
172
+ <section id="ask">
173
+ <h2>Ask bankML</h2>
174
+ <p class="mantra" id="mantra"></p>
175
+ <div class="modes" role="radiogroup" aria-label="who answers">
176
+ <label><input type="radio" name="mode" value="local" checked> <b>Your own bankML</b> — verified, with a receipt, on your CPU (free)</label>
177
+ <label><input type="radio" name="mode" value="hf"> <b>A Hugging Face provider</b> — not bankML, no receipt, your inference quota</label>
178
+ </div>
179
+ <p class="modenote" id="modenote"></p>
180
+ <div id="localrow" class="row">
181
+ <input id="endpoint" aria-label="your bankml serve address" spellcheck="false">
182
+ <button id="connect" type="button">connect</button>
183
+ <span id="localstatus" class="status"></span>
184
+ <details>
185
+ <summary>Start your own bankML for this page (Linux, AVX2; free)</summary>
186
+ <ol>
187
+ <li><code>git clone https://github.com/cryptoAGI/bankml &amp;&amp; cd bankml &amp;&amp; ./install.sh</code> — builds bankML and imports and verifies Bonsai-8B.</li>
188
+ <li><code>./install.sh stop</code> (the installer's own serve does not allow web pages), then:<br>
189
+ <code>target/release/bankml serve .models/Bonsai-8B-Q1_0.gguf --fork ~/.local/share/bankml/forks/Bonsai-8B-Q1_0.gguf.FORK.json --native --allow-origin <span id="origin"></span></code></li>
190
+ <li>Press <b>connect</b>. Your browser may ask to let this page reach your machine: that is the request for your own bankML on 127.0.0.1, and nothing else.</li>
191
+ </ol>
192
+ </details>
193
+ </div>
194
+ <div id="hfrow" class="row" hidden>
195
+ <input id="model" aria-label="provider model (owner/name)" spellcheck="false">
196
+ <button id="signin" type="button">Sign in with Hugging Face</button>
197
+ <button id="signout" type="button" hidden>sign out</button>
198
+ <span id="who" class="status"></span>
199
+ </div>
200
+ <div id="chat" aria-live="polite"></div>
201
+ <div class="row ask">
202
+ <textarea id="message" rows="2" placeholder="Ask bankML about itself, its exactness, its speed… (Ctrl+Enter)"></textarea>
203
+ <button id="send" type="button">Ask</button>
204
+ </div>
205
+ <p id="askstatus" class="status"></p>
206
+ </section>
207
+
208
  <section>
209
  <h2>bankML in one paragraph</h2>
210
  <p>bankML runs language models on the computer you already have: a laptop, a small server, no graphics card
 
478
  mount(layer);
479
  })();
480
  </script>
481
+ <script type="module" src="bankml-chat.js"></script>
482
  </body>
483
  </html>
sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg and b/sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg differ
 
sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg and b/sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg differ
 
sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg and b/sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg differ
 
sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg and b/sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg differ
 
sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg and b/sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg differ
 
sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg and b/sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg differ
 
sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg and b/sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg differ
 
sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg and b/sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg differ
 
sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg and b/sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg differ
 
sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg and b/sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg differ
 
sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg and b/sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg differ
 
sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg and b/sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg differ
 
sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg and b/sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg differ
 
sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg and b/sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg differ
 
sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg and b/sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg differ
 
sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg and b/sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg differ
 
sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg and b/sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg differ
 
sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg and b/sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg differ
 
sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg and b/sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg differ
 
sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg and b/sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg differ
 
sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg and b/sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg differ
 
sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg and b/sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg differ
 
sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg and b/sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg differ
 
sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg and b/sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg differ
 
sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg and b/sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg differ
 
sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg and b/sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg differ
 
sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg and b/sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg differ
 
sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg and b/sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg differ
 
sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg and b/sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg differ
 
sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg and b/sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg differ
 
sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg and b/sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg differ
 
sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg and b/sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg differ
 
sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg and b/sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg differ
 
sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg and b/sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg differ
 
sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg and b/sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg differ
 
sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg CHANGED
Binary files a/sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg and b/sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg differ