Ask bankML: your own bankML (verified, receipt checked in the page) or a Hugging Face provider (labelled not bankML); source at github main cfe81e1 (serve --allow-origin)
Browse filesThis view is limited to 50 files because it contains too many changes. See raw diff
- CHANGELOG.md +11 -1
- LICENSING.md +1 -1
- README.md +12 -5
- bankML/main.rs +2 -1
- bankML/ollama.rs +1 -1
- bankML/serve.rs +51 -5
- bankml-chat.js +204 -0
- docs/BUILD_HISTORY.md +1 -1
- docs/index.html +4 -0
- docs/install.md +1 -0
- docs/modules/gguf.md +1 -1
- docs/thesis.md +106 -2
- docs/usage.md +3 -0
- index.html +59 -0
- sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg +0 -0
- sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg +0 -0
- sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg +0 -0
- sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg +0 -0
- sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg +0 -0
- sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg +0 -0
- sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg +0 -0
- sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg +0 -0
- sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg +0 -0
- sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg +0 -0
- sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg +0 -0
- sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg +0 -0
- sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg +0 -0
- sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg +0 -0
- sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg +0 -0
- sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg +0 -0
- sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg +0 -0
- sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg +0 -0
- sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg +0 -0
- sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg +0 -0
- sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg +0 -0
- sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg +0 -0
- sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg +0 -0
- sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg +0 -0
- sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg +0 -0
- sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg +0 -0
- sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg +0 -0
- sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg +0 -0
- sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg +0 -0
- sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg +0 -0
- sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg +0 -0
- sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg +0 -0
- sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg +0 -0
- sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg +0 -0
- sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg +0 -0
- sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg +0 -0
CHANGELOG.md
CHANGED
|
@@ -27,6 +27,15 @@ decode speed, the third, is measured on an idle machine next.
|
|
| 27 |
against 38.8 ms** (13×), p90 49.3 against 80.6 ms, over the oracle's 1,645 masks. The oracle now computes every
|
| 28 |
mask both ways: **196 / 196 runs, 1,645 / 1,645 masks** identical to llama.cpp b11192 by each.
|
| 29 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
### The code audit and the documentation (2026-10-06)
|
| 31 |
- **Comments, professional and short.** Every source file's comments were cut to the contract, the invariants, the
|
| 32 |
safety reasoning and the exact upstream behaviour that keeps the bits; the history and rationale moved to its
|
|
@@ -155,7 +164,8 @@ measurements.
|
|
| 155 |
× 4 prompts: each sampler alone, typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and
|
| 156 |
a disabling threshold, dynamic temperature at 0, DRY with defaults, custom breakers and with the repeat penalty, all
|
| 157 |
five at once): mindx-gen39 **76 / 76** answers token-identical (3,576 tokens), Bonsai-1.7B **76 / 76** (2,587
|
| 158 |
-
tokens)
|
|
|
|
| 159 |
- `oracle_std_sort` (`testing/sort_oracle.cpp`, libstdc++'s own `std::sort`): **876 / 876** orders identical — sizes 0
|
| 160 |
to 1,000, heavy ties, sorted, reversed and equal keys — the order typical-p's unstable sort leaves equal scores in.
|
| 161 |
|
|
|
|
| 27 |
against 38.8 ms** (13×), p90 49.3 against 80.6 ms, over the oracle's 1,645 masks. The oracle now computes every
|
| 28 |
mask both ways: **196 / 196 runs, 1,645 / 1,645 masks** identical to llama.cpp b11192 by each.
|
| 29 |
|
| 30 |
+
### Your own bankML from a web page (`--allow-origin`)
|
| 31 |
+
- `bankml serve … --allow-origin ORIGIN` lets one named web page call the gateway from a browser: its CORS preflight
|
| 32 |
+
is answered (with Chrome's private-network grant, as a public page reaching a loopback address requires) and its
|
| 33 |
+
answers, streamed or not, carry `Access-Control-Allow-Origin`; any other origin gets none and a 403 preflight. The
|
| 34 |
+
loopback `Host` and JSON-POST rules are unchanged. It is what lets the bankML Space's page
|
| 35 |
+
([PYTHAI/bankml](https://huggingface.co/spaces/PYTHAI/bankml)) talk to a visitor's own bankML, free, on their own
|
| 36 |
+
CPU — the way Savante's page reaches a local engine. Checked live: the preflight, plain and streamed answers with
|
| 37 |
+
the receipt, and another origin refused.
|
| 38 |
+
|
| 39 |
### The code audit and the documentation (2026-10-06)
|
| 40 |
- **Comments, professional and short.** Every source file's comments were cut to the contract, the invariants, the
|
| 41 |
safety reasoning and the exact upstream behaviour that keeps the bits; the history and rationale moved to its
|
|
|
|
| 164 |
× 4 prompts: each sampler alone, typical-p before top-p and min-p's unsorted path, XTC with a clamped probability and
|
| 165 |
a disabling threshold, dynamic temperature at 0, DRY with defaults, custom breakers and with the repeat penalty, all
|
| 166 |
five at once): mindx-gen39 **76 / 76** answers token-identical (3,576 tokens), Bonsai-1.7B **76 / 76** (2,587
|
| 167 |
+
tokens), Bonsai-8B **76 / 76** (2,361 tokens, `oracle_samplers_8b`); **16 / 16** refusals on each, with
|
| 168 |
+
llama-server's message.
|
| 169 |
- `oracle_std_sort` (`testing/sort_oracle.cpp`, libstdc++'s own `std::sort`): **876 / 876** orders identical — sizes 0
|
| 170 |
to 1,000, heavy ties, sorted, reversed and equal keys — the order typical-p's unstable sort leaves equal scores in.
|
| 171 |
|
LICENSING.md
CHANGED
|
@@ -34,7 +34,7 @@ Rules that keep the layers honest:
|
|
| 34 |
| what | licence |
|
| 35 |
|---|---|
|
| 36 |
| `upstream/` (the AVX2 `Q2_0` kernel prepared for llama.cpp) | `MIT`, llama.cpp's licence, so it can be contributed as is |
|
| 37 |
-
| `testing/gguf_guard.py` | vendored from minaiml (same authors); `MIT OR Apache-2.0` here |
|
| 38 |
| `sAGI/voice/knobs/savante_knobs.js` | a build of DreamKnob and React (MIT); regenerate with `node sAGI/voice/knobs/build.mjs` |
|
| 39 |
| `sAGI/voice/cache/`, `sAGI/voice/export/` | Savante's voice, rendered by bankml with Piper and the `en_GB-cori-high` voice (trained on public-domain LibriVox recordings); offered under `MIT OR Apache-2.0` |
|
| 40 |
| models | not part of this repository; bankml imports only models with open-source licences and pins each by sha256 (`sAGI/models.py`) |
|
|
|
|
| 34 |
| what | licence |
|
| 35 |
|---|---|
|
| 36 |
| `upstream/` (the AVX2 `Q2_0` kernel prepared for llama.cpp) | `MIT`, llama.cpp's licence, so it can be contributed as is |
|
| 37 |
+
| `testing/gguf_guard.py` | vendored from [minaiml](https://github.com/minaiml) (same authors); `MIT OR Apache-2.0` here |
|
| 38 |
| `sAGI/voice/knobs/savante_knobs.js` | a build of DreamKnob and React (MIT); regenerate with `node sAGI/voice/knobs/build.mjs` |
|
| 39 |
| `sAGI/voice/cache/`, `sAGI/voice/export/` | Savante's voice, rendered by bankml with Piper and the `en_GB-cori-high` voice (trained on public-domain LibriVox recordings); offered under `MIT OR Apache-2.0` |
|
| 40 |
| models | not part of this repository; bankml imports only models with open-source licences and pins each by sha256 (`sAGI/models.py`) |
|
README.md
CHANGED
|
@@ -5,6 +5,10 @@ colorFrom: green
|
|
| 5 |
colorTo: indigo
|
| 6 |
sdk: static
|
| 7 |
app_file: index.html
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
pinned: true
|
| 9 |
license: mit
|
| 10 |
short_description: Verified 1-bit and ternary LLMs on a CPU, in Rust
|
|
@@ -18,10 +22,12 @@ tags:
|
|
| 18 |
- verified-inference
|
| 19 |
---
|
| 20 |
|
| 21 |
-
> **
|
| 22 |
-
>
|
| 23 |
-
>
|
| 24 |
-
>
|
|
|
|
|
|
|
| 25 |
> [github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml) · licence `MIT OR Apache-2.0`.
|
| 26 |
|
| 27 |
<h1 align="center">bankML</h1>
|
|
@@ -48,7 +54,8 @@ tags:
|
|
| 48 |
<p align="center">
|
| 49 |
<a href="https://deltaverse.pythai.net/bankml"><b>Why bankML</b></a> (the short version, on the web) ·
|
| 50 |
<a href="docs/thesis.md"><b>the thesis</b></a> · <a href="docs/usage.md"><b>usage</b></a> ·
|
| 51 |
-
<a href="CHANGELOG.md"><b>changelog</b></a>
|
|
|
|
| 52 |
</p>
|
| 53 |
|
| 54 |
---
|
|
|
|
| 5 |
colorTo: indigo
|
| 6 |
sdk: static
|
| 7 |
app_file: index.html
|
| 8 |
+
hf_oauth: true
|
| 9 |
+
hf_oauth_expiration_minutes: 480
|
| 10 |
+
hf_oauth_scopes:
|
| 11 |
+
- inference-api
|
| 12 |
pinned: true
|
| 13 |
license: mit
|
| 14 |
short_description: Verified 1-bit and ternary LLMs on a CPU, in Rust
|
|
|
|
| 22 |
- verified-inference
|
| 23 |
---
|
| 24 |
|
| 25 |
+
> **Ask bankML, two ways, never confused.** *Your own bankML*: this page talks to `bankml serve` on your machine
|
| 26 |
+
> (started with `--allow-origin https://pythai-bankml.static.hf.space`) — bankML's verified arithmetic on your CPU,
|
| 27 |
+
> free, with a receipt whose sha256 the page checks. *A Hugging Face provider*: sign in, and a provider-hosted model
|
| 28 |
+
> answers with bankML's persona (`sAGI/personas/bankml.persona`) on your inference quota — labelled on every answer
|
| 29 |
+
> as not bankML, with no receipt. This Space also holds bankML's whole source, and is ready to run the engine itself
|
| 30 |
+
> (`Dockerfile`, `hf/start.sh`) once it has Docker hardware. Source of record:
|
| 31 |
> [github.com/cryptoAGI/bankml](https://github.com/cryptoAGI/bankml) · licence `MIT OR Apache-2.0`.
|
| 32 |
|
| 33 |
<h1 align="center">bankML</h1>
|
|
|
|
| 54 |
<p align="center">
|
| 55 |
<a href="https://deltaverse.pythai.net/bankml"><b>Why bankML</b></a> (the short version, on the web) ·
|
| 56 |
<a href="docs/thesis.md"><b>the thesis</b></a> · <a href="docs/usage.md"><b>usage</b></a> ·
|
| 57 |
+
<a href="CHANGELOG.md"><b>changelog</b></a> ·
|
| 58 |
+
<a href="https://huggingface.co/spaces/PYTHAI/bankml"><b>on Hugging Face</b></a>
|
| 59 |
</p>
|
| 60 |
|
| 61 |
---
|
bankML/main.rs
CHANGED
|
@@ -13,7 +13,7 @@ const USAGE: &str = "usage: bankml usage [PID …]
|
|
| 13 |
bankml pin FILE --fork FORK.json
|
| 14 |
bankml verify FILE --fork FORK.json [--engine mainline|prism] [--json]
|
| 15 |
bankml serve FILE --fork FORK.json [--upstream HOST:PORT | --spawn LLAMA_SERVER] [--listen HOST:PORT] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR]
|
| 16 |
-
bankml serve FILE --fork FORK.json --native [--listen HOST:PORT] [--upstream HOST:PORT] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR]
|
| 17 |
(answers from bankML's own forward pass; also serves the engine address;
|
| 18 |
OpenAI /v1 and Ollama /api; --registry: every model pinned in DIR,
|
| 19 |
default ~/.local/share/bankml/forks, by name, one resident at a time)
|
|
@@ -293,6 +293,7 @@ fn main() {
|
|
| 293 |
// `--registry` without DIR: the importer's forks directory.
|
| 294 |
registry: a.iter().any(|x| x == "--registry").then(|| registry_dir(&a)),
|
| 295 |
keep_alive: opt("--keep-alive"),
|
|
|
|
| 296 |
};
|
| 297 |
match bankml::serve::run(cfg) {
|
| 298 |
Ok(()) => 0,
|
|
|
|
| 13 |
bankml pin FILE --fork FORK.json
|
| 14 |
bankml verify FILE --fork FORK.json [--engine mainline|prism] [--json]
|
| 15 |
bankml serve FILE --fork FORK.json [--upstream HOST:PORT | --spawn LLAMA_SERVER] [--listen HOST:PORT] [--threads N] [--ctx N] [--spec-ngram] [--slot-dir DIR]
|
| 16 |
+
bankml serve FILE --fork FORK.json --native [--listen HOST:PORT] [--upstream HOST:PORT] [--ctx N] [--registry [DIR]] [--keep-alive DUR] [--slot-dir DIR] [--allow-origin ORIGIN]
|
| 17 |
(answers from bankML's own forward pass; also serves the engine address;
|
| 18 |
OpenAI /v1 and Ollama /api; --registry: every model pinned in DIR,
|
| 19 |
default ~/.local/share/bankml/forks, by name, one resident at a time)
|
|
|
|
| 293 |
// `--registry` without DIR: the importer's forks directory.
|
| 294 |
registry: a.iter().any(|x| x == "--registry").then(|| registry_dir(&a)),
|
| 295 |
keep_alive: opt("--keep-alive"),
|
| 296 |
+
allow_origin: opt("--allow-origin"),
|
| 297 |
};
|
| 298 |
match bankml::serve::run(cfg) {
|
| 299 |
Ok(()) => 0,
|
bankML/ollama.rs
CHANGED
|
@@ -522,7 +522,7 @@ fn answer(c: &mut TcpStream, l: &Loaded, req: &Json, msgs: Option<&Json>, o: Opt
|
|
| 522 |
};
|
| 523 |
let mut content = crate::grammar::ContentStream::new(&constraint);
|
| 524 |
if stream {
|
| 525 |
-
write!(c, "HTTP/1.1 200 OK\r\nContent-Type: application/x-ndjson\r\nCache-Control: no-cache\r\
|
| 526 |
let send = |c: &mut TcpStream, piece: &str| piece.is_empty() || c.write_all(piece_line(model, chat, piece).as_bytes()).and_then(|_| c.flush()).is_ok();
|
| 527 |
let done = eng.complete(&prompt, params, o.max, grammar, |piece| {
|
| 528 |
if t.ttft.is_none() {
|
|
|
|
| 522 |
};
|
| 523 |
let mut content = crate::grammar::ContentStream::new(&constraint);
|
| 524 |
if stream {
|
| 525 |
+
write!(c, "HTTP/1.1 200 OK\r\nContent-Type: application/x-ndjson\r\nCache-Control: no-cache\r\n{}Connection: close\r\n\r\n", crate::serve::cors_headers())?;
|
| 526 |
let send = |c: &mut TcpStream, piece: &str| piece.is_empty() || c.write_all(piece_line(model, chat, piece).as_bytes()).and_then(|_| c.flush()).is_ok();
|
| 527 |
let done = eng.complete(&prompt, params, o.max, grammar, |piece| {
|
| 528 |
if t.ttft.is_none() {
|
bankML/serve.rs
CHANGED
|
@@ -43,6 +43,25 @@ pub struct Config {
|
|
| 43 |
pub registry: Option<PathBuf>,
|
| 44 |
/// `--native`: how long `/api/*` keeps a model resident when the request does not say (Ollama's default, 5m)
|
| 45 |
pub keep_alive: Option<String>,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 46 |
}
|
| 47 |
|
| 48 |
struct State {
|
|
@@ -57,6 +76,8 @@ struct State {
|
|
| 57 |
ident: FileIdent,
|
| 58 |
/// where `/slots/{id}?action=save|restore` keeps slot files (`--slot-dir`; native mode)
|
| 59 |
slot_dir: Option<PathBuf>,
|
|
|
|
|
|
|
| 60 |
}
|
| 61 |
|
| 62 |
/// The file identity checked before each P0 answer: device, inode, size and mtime, not the bytes (re-hashing per
|
|
@@ -110,6 +131,9 @@ mod sig {
|
|
| 110 |
|
| 111 |
/// Verify the model, launch or check the upstream (or start `--native`), then serve until killed.
|
| 112 |
pub fn run(cfg: Config) -> Result<(), String> {
|
|
|
|
|
|
|
|
|
|
| 113 |
let before = ident(&cfg.model).map_err(|e| format!("{}: {e}", cfg.model.display()))?;
|
| 114 |
let verified = crate::verify(&cfg.model, &cfg.fork_json, cfg.engine)?;
|
| 115 |
let model = cfg.model.canonicalize().map_err(|e| format!("{}: {e}", cfg.model.display()))?;
|
|
@@ -153,7 +177,7 @@ pub fn run(cfg: Config) -> Result<(), String> {
|
|
| 153 |
}
|
| 154 |
let hashed_at = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
|
| 155 |
let engine = "llama.cpp b11192 llama-server (loopback), behind bankml P0".to_string();
|
| 156 |
-
let st = Arc::new(State { native: None, keep_alive: crate::native::KeepAlive::Forever, verified, model, upstream, engine, hashed_at, ident: id, slot_dir: None });
|
| 157 |
let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
|
| 158 |
eprintln!("bankml serve {}: {} verified (sha256 {}), upstream {} serves it; listening on http://{}",
|
| 159 |
crate::VERSION, st.model.display(), st.verified.model_sha256, st.upstream, cfg.listen);
|
|
@@ -189,7 +213,7 @@ fn run_native(cfg: Config, verified: Verified, model: PathBuf, id: FileIdent, up
|
|
| 189 |
if let Some(d) = &cfg.slot_dir {
|
| 190 |
std::fs::create_dir_all(d).map_err(|e| format!("--slot-dir {}: {e}", d.display()))?;
|
| 191 |
}
|
| 192 |
-
let st = Arc::new(State { slot_dir: cfg.slot_dir.clone(), native: Some(rs), keep_alive: ka, verified: Verified { model_sha256: sha.clone(), guard: "play", engine: cfg.engine.as_str(), arch: None, name: None, types: Vec::new() },
|
| 193 |
model, upstream: upstream.clone(), engine, hashed_at, ident: id });
|
| 194 |
let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
|
| 195 |
let lu = TcpListener::bind(&upstream).map_err(|e| format!("cannot listen on the engine address {upstream}: {e} (is llama-server running there?)"))?;
|
|
@@ -378,7 +402,7 @@ pub(crate) fn respond(c: &mut TcpStream, code: u16, ctype: &str, body: &[u8]) ->
|
|
| 378 |
_ => "Error",
|
| 379 |
};
|
| 380 |
let code = if (100..600).contains(&code) { code } else { 502 }; // a garbled upstream head is a bad gateway
|
| 381 |
-
write!(c, "HTTP/1.1 {code} {reason}\r\nContent-Type: {ctype}\r\nContent-Length: {}\r\
|
| 382 |
c.write_all(body)
|
| 383 |
}
|
| 384 |
|
|
@@ -405,6 +429,18 @@ fn handle(mut c: TcpStream, st: &State) -> std::io::Result<()> {
|
|
| 405 |
if !loopback_host(header(&h, "host").unwrap_or("")) {
|
| 406 |
return refuse(403, b"bankml serve answers loopback clients only (Host must be 127.0.0.1, localhost or [::1])", &mut r);
|
| 407 |
}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 408 |
if header(&h, "transfer-encoding").is_some() {
|
| 409 |
return refuse(400, b"chunked requests are not accepted; send Content-Length", &mut r);
|
| 410 |
}
|
|
@@ -583,7 +619,7 @@ fn native_chat(c: &mut TcpStream, rs: &crate::native::Residency, body: &[u8]) ->
|
|
| 583 |
let model_id = l.model.file_name().map(|n| n.to_string_lossy().into_owned()).unwrap_or_default();
|
| 584 |
let created = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
|
| 585 |
let r = if stream {
|
| 586 |
-
write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\
|
| 587 |
// with logprobs, llama-server's shape: the first token's partial opens with the role delta, and each token's
|
| 588 |
// entry rides on the last delta its partial sends (no delta, no entry)
|
| 589 |
let logprobs = nc.params.n_probs > 0;
|
|
@@ -978,7 +1014,7 @@ fn chat(c: &mut TcpStream, st: &State, body: &[u8]) -> std::io::Result<()> {
|
|
| 978 |
let merged = format!("{}, \"bankml_receipt\": {}}}", &s[..end], t.receipt(&st.engine, &st.verified));
|
| 979 |
return respond(c, 200, "application/json", merged.as_bytes());
|
| 980 |
}
|
| 981 |
-
write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\
|
| 982 |
let mut pending = Vec::new();
|
| 983 |
while let Some(ch) = b.next(&mut r)? {
|
| 984 |
pending.extend(ch);
|
|
@@ -1268,6 +1304,16 @@ mod tests {
|
|
| 1268 |
assert_eq!(header(&h, "host"), Some("127.0.0.1:18093"));
|
| 1269 |
}
|
| 1270 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1271 |
#[test]
|
| 1272 |
fn only_loopback_hosts() {
|
| 1273 |
for h in ["127.0.0.1", "127.0.0.1:18093", "localhost:18093", "LOCALHOST", "[::1]:18093", "[::1]"] {
|
|
|
|
| 43 |
pub registry: Option<PathBuf>,
|
| 44 |
/// `--native`: how long `/api/*` keeps a model resident when the request does not say (Ollama's default, 5m)
|
| 45 |
pub keep_alive: Option<String>,
|
| 46 |
+
/// one web page origin (`https://host[:port]`) whose scripts may call this gateway from a browser: CORS and
|
| 47 |
+
/// Chrome's private-network preflight are answered for it alone; the loopback `Host` rule still holds
|
| 48 |
+
pub allow_origin: Option<String>,
|
| 49 |
+
}
|
| 50 |
+
|
| 51 |
+
/// Whether `o` is an origin: `http(s)://host[:port]`, no path, query or credentials.
|
| 52 |
+
pub fn is_origin(o: &str) -> bool {
|
| 53 |
+
let rest = o.strip_prefix("https://").or_else(|| o.strip_prefix("http://"));
|
| 54 |
+
rest.is_some_and(|r| !r.is_empty() && r.chars().all(|c| c.is_ascii_alphanumeric() || matches!(c, '.' | '-' | ':' | '[' | ']')))
|
| 55 |
+
}
|
| 56 |
+
|
| 57 |
+
thread_local! {
|
| 58 |
+
/// The CORS headers for the request this connection thread is answering (empty unless its `Origin` is allowed).
|
| 59 |
+
static CORS: std::cell::RefCell<String> = const { std::cell::RefCell::new(String::new()) };
|
| 60 |
+
}
|
| 61 |
+
|
| 62 |
+
/// The CORS header lines to add to this connection's response (`\r\n`-terminated; empty when none apply).
|
| 63 |
+
pub(crate) fn cors_headers() -> String {
|
| 64 |
+
CORS.with(|c| c.borrow().clone())
|
| 65 |
}
|
| 66 |
|
| 67 |
struct State {
|
|
|
|
| 76 |
ident: FileIdent,
|
| 77 |
/// where `/slots/{id}?action=save|restore` keeps slot files (`--slot-dir`; native mode)
|
| 78 |
slot_dir: Option<PathBuf>,
|
| 79 |
+
/// `--allow-origin`: the one web origin answered with CORS headers
|
| 80 |
+
allow_origin: Option<String>,
|
| 81 |
}
|
| 82 |
|
| 83 |
/// The file identity checked before each P0 answer: device, inode, size and mtime, not the bytes (re-hashing per
|
|
|
|
| 131 |
|
| 132 |
/// Verify the model, launch or check the upstream (or start `--native`), then serve until killed.
|
| 133 |
pub fn run(cfg: Config) -> Result<(), String> {
|
| 134 |
+
if let Some(o) = cfg.allow_origin.as_deref().filter(|o| !is_origin(o)) {
|
| 135 |
+
return Err(format!("--allow-origin {o}: an origin is http(s)://host[:port], with no path"));
|
| 136 |
+
}
|
| 137 |
let before = ident(&cfg.model).map_err(|e| format!("{}: {e}", cfg.model.display()))?;
|
| 138 |
let verified = crate::verify(&cfg.model, &cfg.fork_json, cfg.engine)?;
|
| 139 |
let model = cfg.model.canonicalize().map_err(|e| format!("{}: {e}", cfg.model.display()))?;
|
|
|
|
| 177 |
}
|
| 178 |
let hashed_at = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
|
| 179 |
let engine = "llama.cpp b11192 llama-server (loopback), behind bankml P0".to_string();
|
| 180 |
+
let st = Arc::new(State { native: None, keep_alive: crate::native::KeepAlive::Forever, verified, model, upstream, engine, hashed_at, ident: id, slot_dir: None, allow_origin: cfg.allow_origin.clone() });
|
| 181 |
let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
|
| 182 |
eprintln!("bankml serve {}: {} verified (sha256 {}), upstream {} serves it; listening on http://{}",
|
| 183 |
crate::VERSION, st.model.display(), st.verified.model_sha256, st.upstream, cfg.listen);
|
|
|
|
| 213 |
if let Some(d) = &cfg.slot_dir {
|
| 214 |
std::fs::create_dir_all(d).map_err(|e| format!("--slot-dir {}: {e}", d.display()))?;
|
| 215 |
}
|
| 216 |
+
let st = Arc::new(State { slot_dir: cfg.slot_dir.clone(), allow_origin: cfg.allow_origin.clone(), native: Some(rs), keep_alive: ka, verified: Verified { model_sha256: sha.clone(), guard: "play", engine: cfg.engine.as_str(), arch: None, name: None, types: Vec::new() },
|
| 217 |
model, upstream: upstream.clone(), engine, hashed_at, ident: id });
|
| 218 |
let l = TcpListener::bind(&cfg.listen).map_err(|e| format!("cannot listen on {}: {e}", cfg.listen))?;
|
| 219 |
let lu = TcpListener::bind(&upstream).map_err(|e| format!("cannot listen on the engine address {upstream}: {e} (is llama-server running there?)"))?;
|
|
|
|
| 402 |
_ => "Error",
|
| 403 |
};
|
| 404 |
let code = if (100..600).contains(&code) { code } else { 502 }; // a garbled upstream head is a bad gateway
|
| 405 |
+
write!(c, "HTTP/1.1 {code} {reason}\r\nContent-Type: {ctype}\r\nContent-Length: {}\r\n{}Connection: close\r\n\r\n", body.len(), cors_headers())?;
|
| 406 |
c.write_all(body)
|
| 407 |
}
|
| 408 |
|
|
|
|
| 429 |
if !loopback_host(header(&h, "host").unwrap_or("")) {
|
| 430 |
return refuse(403, b"bankml serve answers loopback clients only (Host must be 127.0.0.1, localhost or [::1])", &mut r);
|
| 431 |
}
|
| 432 |
+
// --allow-origin: that one web page may read the answers from a browser; every other origin gets no CORS headers
|
| 433 |
+
let allowed = header(&h, "origin").filter(|o| st.allow_origin.as_deref() == Some(*o)).map(str::to_string);
|
| 434 |
+
CORS.with(|c| *c.borrow_mut() = allowed.as_deref().map(|o| format!("Access-Control-Allow-Origin: {o}\r\nVary: Origin\r\n")).unwrap_or_default());
|
| 435 |
+
if method == "OPTIONS" {
|
| 436 |
+
if allowed.is_none() {
|
| 437 |
+
return refuse(403, b"bankml serve: this origin is not allowed (start serve with --allow-origin ORIGIN)", &mut r);
|
| 438 |
+
}
|
| 439 |
+
// the preflight; Chrome asks before a public page may reach a loopback address (Private Network Access)
|
| 440 |
+
let pna = header(&h, "access-control-request-private-network").is_some_and(|v| v.eq_ignore_ascii_case("true"));
|
| 441 |
+
return write!(c, "HTTP/1.1 204 No Content\r\n{}Access-Control-Allow-Methods: GET, POST\r\nAccess-Control-Allow-Headers: Content-Type\r\n{}Access-Control-Max-Age: 600\r\nContent-Length: 0\r\nConnection: close\r\n\r\n",
|
| 442 |
+
cors_headers(), if pna { "Access-Control-Allow-Private-Network: true\r\n" } else { "" });
|
| 443 |
+
}
|
| 444 |
if header(&h, "transfer-encoding").is_some() {
|
| 445 |
return refuse(400, b"chunked requests are not accepted; send Content-Length", &mut r);
|
| 446 |
}
|
|
|
|
| 619 |
let model_id = l.model.file_name().map(|n| n.to_string_lossy().into_owned()).unwrap_or_default();
|
| 620 |
let created = SystemTime::now().duration_since(UNIX_EPOCH).map(|d| d.as_secs()).unwrap_or(0);
|
| 621 |
let r = if stream {
|
| 622 |
+
write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\n{}Connection: close\r\n\r\n", cors_headers())?;
|
| 623 |
// with logprobs, llama-server's shape: the first token's partial opens with the role delta, and each token's
|
| 624 |
// entry rides on the last delta its partial sends (no delta, no entry)
|
| 625 |
let logprobs = nc.params.n_probs > 0;
|
|
|
|
| 1014 |
let merged = format!("{}, \"bankml_receipt\": {}}}", &s[..end], t.receipt(&st.engine, &st.verified));
|
| 1015 |
return respond(c, 200, "application/json", merged.as_bytes());
|
| 1016 |
}
|
| 1017 |
+
write!(c, "HTTP/1.1 200 OK\r\nContent-Type: text/event-stream\r\nCache-Control: no-cache\r\n{}Connection: close\r\n\r\n", cors_headers())?;
|
| 1018 |
let mut pending = Vec::new();
|
| 1019 |
while let Some(ch) = b.next(&mut r)? {
|
| 1020 |
pending.extend(ch);
|
|
|
|
| 1304 |
assert_eq!(header(&h, "host"), Some("127.0.0.1:18093"));
|
| 1305 |
}
|
| 1306 |
|
| 1307 |
+
#[test]
|
| 1308 |
+
fn origins_are_scheme_host_port_only() {
|
| 1309 |
+
for o in ["https://pythai-bankml.static.hf.space", "http://127.0.0.1:8000", "https://[::1]:7860"] {
|
| 1310 |
+
assert!(is_origin(o), "{o}");
|
| 1311 |
+
}
|
| 1312 |
+
for o in ["https://a.example/path", "pythai-bankml.static.hf.space", "https://", "https://u@h", "https://h?x=1", "*"] {
|
| 1313 |
+
assert!(!is_origin(o), "{o}");
|
| 1314 |
+
}
|
| 1315 |
+
}
|
| 1316 |
+
|
| 1317 |
#[test]
|
| 1318 |
fn only_loopback_hosts() {
|
| 1319 |
for h in ["127.0.0.1", "127.0.0.1:18093", "localhost:18093", "LOCALHOST", "[::1]:18093", "[::1]"] {
|
bankml-chat.js
ADDED
|
@@ -0,0 +1,204 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
// SPDX-License-Identifier: MIT OR Apache-2.0
|
| 2 |
+
// Ask bankML, from this static page: two ways, never confused.
|
| 3 |
+
// · Your own bankML — this browser talks to `bankml serve` on the visitor's machine (started with
|
| 4 |
+
// --allow-origin for this page): bankML's verified arithmetic, the visitor's CPU and RAM, a receipt on every
|
| 5 |
+
// answer whose sha256 this page checks. Free.
|
| 6 |
+
// · A Hugging Face provider — sign in with Hugging Face; a provider-hosted model answers under bankML's persona,
|
| 7 |
+
// billed to the visitor's inference quota. NOT bankML's arithmetic, no receipt; said on every answer.
|
| 8 |
+
// Nothing an engine or a provider returns becomes markup: text only.
|
| 9 |
+
import { oauthLoginUrl, oauthHandleRedirectIfPresent } from "https://cdn.jsdelivr.net/npm/@huggingface/hub@2.17.1/+esm";
|
| 10 |
+
import { InferenceClient } from "https://cdn.jsdelivr.net/npm/@huggingface/inference@4.13.28/+esm";
|
| 11 |
+
|
| 12 |
+
const $ = (id) => document.getElementById(id);
|
| 13 |
+
const SESSION = "bankml.oauth";
|
| 14 |
+
const DEFAULT_ENDPOINT = "http://127.0.0.1:18093";
|
| 15 |
+
const DEFAULT_MODEL = "Qwen/Qwen3-8B";
|
| 16 |
+
const history = [];
|
| 17 |
+
let persona = null;
|
| 18 |
+
let oauth = null;
|
| 19 |
+
|
| 20 |
+
const mode = () => document.querySelector('input[name="mode"]:checked').value;
|
| 21 |
+
|
| 22 |
+
// ── the persona: the same file the bankML console speaks from ─────────────────────────────────────────────────
|
| 23 |
+
async function loadPersona() {
|
| 24 |
+
try {
|
| 25 |
+
const r = await fetch("sAGI/personas/bankml.persona", { cache: "no-store" });
|
| 26 |
+
persona = await r.json();
|
| 27 |
+
} catch (e) {
|
| 28 |
+
persona = null;
|
| 29 |
+
$("askstatus").textContent = "the persona could not be read: " + e;
|
| 30 |
+
}
|
| 31 |
+
}
|
| 32 |
+
|
| 33 |
+
// ── the SELF block, as the console builds it: one sentence per measurement, "not measured" when it was not ──────
|
| 34 |
+
const get = async (ep, path) => {
|
| 35 |
+
try {
|
| 36 |
+
const r = await fetch(ep + path, { cache: "no-store" });
|
| 37 |
+
return r.ok ? await r.json() : null;
|
| 38 |
+
} catch { return null; }
|
| 39 |
+
};
|
| 40 |
+
async function selfText(ep) {
|
| 41 |
+
const [b, u, m] = [await get(ep, "/bankml") || {}, await get(ep, "/bankml/usage") || {}, await get(ep, "/bankml/metrics") || {}];
|
| 42 |
+
const last = (m.records || [{}]).slice(-1)[0] || {};
|
| 43 |
+
const v = (x, unit = "", scale = 1, d = 1) => (x === null || x === undefined ? "not measured" : (x * scale).toFixed(d) + unit);
|
| 44 |
+
const ver = b.verified || {};
|
| 45 |
+
return [
|
| 46 |
+
`- the model I am running: ${ver.name || b.resident || "none"} (sha256 ${String(ver.model_sha256 || "").slice(0, 16)}…), bankML ${ver.bankml}`,
|
| 47 |
+
`- tokens I have read in total (prompts): ${v(m.prompt_tokens, "", 1, 0)}`,
|
| 48 |
+
`- tokens I have written in total (answers): ${v(m.completion_tokens, "", 1, 0)}`,
|
| 49 |
+
`- time to first token of my last answer: ${v(last.ttft_ms, " milliseconds", 1, 0)}`,
|
| 50 |
+
`- prompt reading speed of my last answer: ${v(last.prompt_tps, " tokens per second")}`,
|
| 51 |
+
`- generation speed of my last answer: ${v(last.eval_tps, " tokens per second")}`,
|
| 52 |
+
`- CPU use right now: ${v(u.cpu_percent, " percent of one core")}`,
|
| 53 |
+
`- memory I hold (RSS): ${v(u.rss_bytes, " GB", 1e-9, 2)}; memory still available on this machine: ${v(u.mem_available_bytes, " GB", 1e-9, 2)}`,
|
| 54 |
+
`- power the CPU package draws: ${v(u.package_watts, " watts")}; energy per token I write: ${v(m.joules_per_token, " joules", 1, 3)}`,
|
| 55 |
+
].join("\n");
|
| 56 |
+
}
|
| 57 |
+
|
| 58 |
+
// ── the conversation ──────────────────────────────────────────────────────────────────────────────────────────
|
| 59 |
+
function bubble(role, text) {
|
| 60 |
+
const d = document.createElement("div");
|
| 61 |
+
d.className = "msg " + role;
|
| 62 |
+
d.textContent = text;
|
| 63 |
+
$("chat").append(d);
|
| 64 |
+
d.scrollIntoView({ block: "nearest" });
|
| 65 |
+
return d;
|
| 66 |
+
}
|
| 67 |
+
function note(el, text, cls) {
|
| 68 |
+
const p = document.createElement("div");
|
| 69 |
+
p.className = "meta " + (cls || "");
|
| 70 |
+
p.textContent = text;
|
| 71 |
+
el.after(p);
|
| 72 |
+
return p;
|
| 73 |
+
}
|
| 74 |
+
async function sha256hex(text) {
|
| 75 |
+
const h = await crypto.subtle.digest("SHA-256", new TextEncoder().encode(text));
|
| 76 |
+
return [...new Uint8Array(h)].map((b) => b.toString(16).padStart(2, "0")).join("");
|
| 77 |
+
}
|
| 78 |
+
|
| 79 |
+
// ── your own bankML ───────────────────────────────────────────────────────────────────────────────────────────
|
| 80 |
+
async function connect() {
|
| 81 |
+
const ep = $("endpoint").value.trim().replace(/\/+$/, "") || DEFAULT_ENDPOINT;
|
| 82 |
+
$("localstatus").textContent = "connecting…";
|
| 83 |
+
const b = await get(ep, "/bankml");
|
| 84 |
+
if (!b) {
|
| 85 |
+
$("localstatus").textContent = "not reachable — is bankml serve running with --allow-origin " + location.origin + " ? (see below)";
|
| 86 |
+
return false;
|
| 87 |
+
}
|
| 88 |
+
const v = b.verified || {};
|
| 89 |
+
$("localstatus").textContent = v.guard === "play"
|
| 90 |
+
? `✓ connected: ${v.name || b.resident || "model"} · bankML ${v.bankml} · sha256 ${String(v.model_sha256 || "").slice(0, 12)}…`
|
| 91 |
+
: "connected, but no verified model is loaded yet";
|
| 92 |
+
return v.guard === "play";
|
| 93 |
+
}
|
| 94 |
+
async function askLocal(message) {
|
| 95 |
+
const ep = $("endpoint").value.trim().replace(/\/+$/, "") || DEFAULT_ENDPOINT;
|
| 96 |
+
const system = persona.system_prompt + "\n\nSELF (measured by bankML just now):\n" + await selfText(ep);
|
| 97 |
+
const out = bubble("assistant", "…");
|
| 98 |
+
let text = "", receipt = null;
|
| 99 |
+
const r = await fetch(ep + "/v1/chat/completions", {
|
| 100 |
+
method: "POST", headers: { "Content-Type": "application/json" },
|
| 101 |
+
body: JSON.stringify({ messages: [{ role: "system", content: system }, ...history.slice(-12), { role: "user", content: message }],
|
| 102 |
+
stream: true, max_tokens: 384 }),
|
| 103 |
+
});
|
| 104 |
+
if (!r.ok) throw new Error(`HTTP ${r.status}: ${(await r.text()).slice(0, 300)}`);
|
| 105 |
+
const reader = r.body.getReader(), dec = new TextDecoder();
|
| 106 |
+
let buf = "";
|
| 107 |
+
for (;;) {
|
| 108 |
+
const { done, value } = await reader.read();
|
| 109 |
+
if (done) break;
|
| 110 |
+
buf += dec.decode(value, { stream: true });
|
| 111 |
+
let i;
|
| 112 |
+
while ((i = buf.indexOf("\n")) >= 0) {
|
| 113 |
+
const line = buf.slice(0, i).trim();
|
| 114 |
+
buf = buf.slice(i + 1);
|
| 115 |
+
if (!line.startsWith("data:") || line === "data: [DONE]") continue;
|
| 116 |
+
const d = JSON.parse(line.slice(5));
|
| 117 |
+
if (d.bankml_receipt) { receipt = d.bankml_receipt; continue; }
|
| 118 |
+
const piece = d.choices?.[0]?.delta?.content;
|
| 119 |
+
if (piece) { text += piece; out.textContent = text; }
|
| 120 |
+
}
|
| 121 |
+
}
|
| 122 |
+
const ok = receipt && (await sha256hex(text)) === receipt.response_sha256;
|
| 123 |
+
note(out, receipt
|
| 124 |
+
? `${ok ? "✓" : "✗"} receipt — bankML ${receipt.bankml} · model sha256 ${String(receipt.model_sha256 || "").slice(0, 12)}… · answer sha256 ${ok ? "matches the text received" : "does NOT match the text received"}`
|
| 125 |
+
: "no receipt came with this answer", ok ? "ok" : "bad");
|
| 126 |
+
history.push({ role: "user", content: message }, { role: "assistant", content: text });
|
| 127 |
+
}
|
| 128 |
+
|
| 129 |
+
// ── a Hugging Face provider (not bankML) ──────────────────────────────────────────────────────────────────────
|
| 130 |
+
function readSession() { try { return JSON.parse(sessionStorage.getItem(SESSION) || "null"); } catch { return null; } }
|
| 131 |
+
function writeSession(v) { try { v ? sessionStorage.setItem(SESSION, JSON.stringify(v)) : sessionStorage.removeItem(SESSION); } catch {} }
|
| 132 |
+
const signedIn = () => !!(oauth && oauth.accessToken && new Date(oauth.accessTokenExpiresAt) > new Date());
|
| 133 |
+
function renderAuth() {
|
| 134 |
+
$("signin").hidden = signedIn();
|
| 135 |
+
$("signout").hidden = !signedIn();
|
| 136 |
+
$("who").textContent = signedIn()
|
| 137 |
+
? `signed in as ${oauth.userInfo?.preferred_username || oauth.userInfo?.name || "you"} — answers spend your inference quota`
|
| 138 |
+
: "not signed in";
|
| 139 |
+
}
|
| 140 |
+
async function askProvider(message) {
|
| 141 |
+
const model = $("model").value.trim() || DEFAULT_MODEL;
|
| 142 |
+
if (!/^[\w.-]+\/[\w.-]+$/.test(model)) throw new Error("the model must be a Hugging Face repository id, owner/name");
|
| 143 |
+
// the persona, told the truth about where it is running
|
| 144 |
+
const system = persona.system_prompt + `\n\nIMPORTANT: this answer is NOT produced by bankML. It is produced by ${model} through a Hugging Face inference provider: there is no bankML arithmetic, no receipt and no SELF block. Say so if you are asked about your speed, your use, your receipt or your verification.`;
|
| 145 |
+
const out = bubble("assistant", "…");
|
| 146 |
+
let text = "";
|
| 147 |
+
const client = new InferenceClient(oauth.accessToken);
|
| 148 |
+
const stream = client.chatCompletionStream({ provider: "auto", model, max_tokens: 384,
|
| 149 |
+
messages: [{ role: "system", content: system }, ...history.slice(-12), { role: "user", content: message }] });
|
| 150 |
+
for await (const chunk of stream) {
|
| 151 |
+
const piece = chunk?.choices?.[0]?.delta?.content;
|
| 152 |
+
if (piece) { text += piece; out.textContent = text; }
|
| 153 |
+
}
|
| 154 |
+
note(out, `not bankML — ${model} via a Hugging Face provider · no receipt · your quota`, "warn");
|
| 155 |
+
history.push({ role: "user", content: message }, { role: "assistant", content: text });
|
| 156 |
+
}
|
| 157 |
+
|
| 158 |
+
// ── wiring ─────────────────────────────────────────────────────────────────────────────────────────────────────
|
| 159 |
+
function renderMode() {
|
| 160 |
+
const m = mode();
|
| 161 |
+
$("localrow").hidden = m !== "local";
|
| 162 |
+
$("hfrow").hidden = m !== "hf";
|
| 163 |
+
$("modenote").textContent = m === "local"
|
| 164 |
+
? "bankML's own arithmetic on your machine: verified model, receipt checked here, nothing sent anywhere else."
|
| 165 |
+
: "Not bankML: a provider-hosted model speaks with bankML's persona, without its arithmetic or a receipt.";
|
| 166 |
+
}
|
| 167 |
+
document.querySelectorAll('input[name="mode"]').forEach((r) => r.addEventListener("change", renderMode));
|
| 168 |
+
$("connect").addEventListener("click", connect);
|
| 169 |
+
$("signin").addEventListener("click", async () => {
|
| 170 |
+
if (!window.huggingface?.variables?.OAUTH_CLIENT_ID) { $("who").textContent = "sign-in works only on the Hugging Face Space itself"; return; }
|
| 171 |
+
window.location.href = await oauthLoginUrl({ scopes: window.huggingface.variables.OAUTH_SCOPES });
|
| 172 |
+
});
|
| 173 |
+
$("signout").addEventListener("click", () => { writeSession(null); oauth = null; renderAuth(); });
|
| 174 |
+
$("send").addEventListener("click", async () => {
|
| 175 |
+
const message = $("message").value.trim();
|
| 176 |
+
if (!message || !persona) return;
|
| 177 |
+
if (mode() === "hf" && !signedIn()) { $("askstatus").textContent = "sign in with Hugging Face first"; return; }
|
| 178 |
+
$("message").value = "";
|
| 179 |
+
$("send").disabled = true;
|
| 180 |
+
$("askstatus").textContent = "";
|
| 181 |
+
bubble("user", message);
|
| 182 |
+
try {
|
| 183 |
+
await (mode() === "local" ? askLocal(message) : askProvider(message));
|
| 184 |
+
} catch (e) {
|
| 185 |
+
const local = mode() === "local";
|
| 186 |
+
$("askstatus").textContent = (local ? "your bankML did not answer: " : "the provider did not answer: ") + String(e.message || e).slice(0, 300)
|
| 187 |
+
+ (local ? " — is bankml serve running with --allow-origin " + location.origin + " ?" : "");
|
| 188 |
+
} finally {
|
| 189 |
+
$("send").disabled = false;
|
| 190 |
+
}
|
| 191 |
+
});
|
| 192 |
+
$("message").addEventListener("keydown", (e) => { if (e.key === "Enter" && (e.ctrlKey || e.metaKey)) $("send").click(); });
|
| 193 |
+
$("origin").textContent = location.origin;
|
| 194 |
+
|
| 195 |
+
(async () => {
|
| 196 |
+
$("endpoint").value = DEFAULT_ENDPOINT;
|
| 197 |
+
$("model").value = DEFAULT_MODEL;
|
| 198 |
+
renderMode();
|
| 199 |
+
await loadPersona();
|
| 200 |
+
if (persona) $("mantra").textContent = "“" + persona.mantra + "”";
|
| 201 |
+
try { const res = await oauthHandleRedirectIfPresent(); if (res) writeSession(res); } catch (e) { $("who").textContent = "sign-in failed: " + String(e).slice(0, 120); }
|
| 202 |
+
oauth = readSession();
|
| 203 |
+
renderAuth();
|
| 204 |
+
})();
|
docs/BUILD_HISTORY.md
CHANGED
|
@@ -107,7 +107,7 @@ then memory traffic per token, with strategies from llama.cpp, vLLM, Ollama and
|
|
| 107 |
dependency justified in `Cargo.toml` comments. A single static binary.
|
| 108 |
3. **Verified response** — every answer carries its evidence, and a model that cannot be
|
| 109 |
verified does not answer:
|
| 110 |
-
- before load: the GGUF guard (port of minaiml `gguf_guard.py`) → `play | refuse | need_more`;
|
| 111 |
refuse is shown with its reason, never a silent fallback (Bonsai 2 Q2_0 loads on mainline
|
| 112 |
and answers in gibberish — the worst failure);
|
| 113 |
- the file's sha256 must match a pinned record (`FORK.json` of the PYTHAI fork);
|
|
|
|
| 107 |
dependency justified in `Cargo.toml` comments. A single static binary.
|
| 108 |
3. **Verified response** — every answer carries its evidence, and a model that cannot be
|
| 109 |
verified does not answer:
|
| 110 |
+
- before load: the GGUF guard (port of [minaiml](https://github.com/minaiml) `gguf_guard.py`) → `play | refuse | need_more`;
|
| 111 |
refuse is shown with its reason, never a silent fallback (Bonsai 2 Q2_0 loads on mainline
|
| 112 |
and answers in gibberish — the worst failure);
|
| 113 |
- the file's sha256 must match a pinned record (`FORK.json` of the PYTHAI fork);
|
docs/index.html
CHANGED
|
@@ -113,6 +113,9 @@ article p[align="center"] { text-align: center; }
|
|
| 113 |
.picker select { width: 100%; font: 15px var(--f-body); padding: 8px; border-radius: 6px; border: 1px solid var(--rule); background: var(--panel); color: var(--ink); }
|
| 114 |
article h1 { font-size: 1.65rem; }
|
| 115 |
}
|
|
|
|
|
|
|
|
|
|
| 116 |
@media (prefers-reduced-motion: no-preference) { html { scroll-behavior: smooth; } }
|
| 117 |
</style>
|
| 118 |
|
|
@@ -120,6 +123,7 @@ article p[align="center"] { text-align: center; }
|
|
| 120 |
<div class="masthead-inner">
|
| 121 |
<p class="brand"><span class="trits" aria-hidden="true"><i></i><i></i><i></i></span>bankml<span class="ver" id="ver">0.2.7</span></p>
|
| 122 |
<span class="tag">Verified low-bit inference for the CPU you already have — the documentation, as of <span id="asof">0.2.7</span>.</span>
|
|
|
|
| 123 |
<div class="proof">gate record 0.3.6: the penalties — repeat, frequency and presence, token-identical to llama-server on <b>56 / 56</b> answers on each of three models, live <b>85 / 85</b> · gate record 0.3.5: JSON schemas — llama-server’s grammar on each template (<b>173 / 173</b> schemas × 3 templates against llama.cpp’s own code) and its answers token-identical on all five native models (<b>28 / 28</b>, <b>11 / 11</b>, <b>56 / 56</b> × 3), the content rule matched on <b>30,063 / 30,063</b> texts per template · <code>bankml create</code>: mindX’s persona layer, made from mindXtrain’s merged output, token-identical end to end (<b>27 / 27</b>, two ways) · gate record 0.3.4: mindX’s own model natively — <code>mindx-gen39</code> and SmolLM2-135M-Instruct (the Llama graph in F16) and Bonsai-1.7B (tied embeddings) token-identical to llama-server b11192: whole model <b>800 / 800</b> and <b>840 / 840</b> rows bit-exact, F16 products <b>552,268 / 552,268</b> elements bit-exact against ggml, seeded sampling <b>40 / 40</b> on each, conversations <b>9 / 9</b>, JSON mode <b>23 / 23</b> · JSON mode — answers under llama-server’s own <code>json_object</code> grammar token-identical, greedy and seeded (<b>23 / 23</b> answers on the 1-bit model, <b>13 / 13</b> on the ternary), and llama.cpp’s grammar sampler matched on <b>1,645 / 1,645</b> whole-vocabulary masks · a C API, <code>libbankml</code>, identical to <code>serve --native</code> · Savante answered by bankML’s own forward pass — whole conversations identical to llama-server on <b>9 / 9</b> turns (text, token counts, prompt-cache reuse) · own forward pass <b>1,064 / 1,064</b> rows bit-exact for the 1-bit and the ternary model · seeded sampling identical on <b>40 / 40</b> continuations · <b>8,188,239,872</b> ternary weights and <b>762 / 762</b> dot products bit-exact against llama.cpp b11192’s compiled library · one ternary token <b>0.222 s vs 2.139 s</b></div>
|
| 124 |
</div>
|
| 125 |
</header>
|
|
|
|
| 113 |
.picker select { width: 100%; font: 15px var(--f-body); padding: 8px; border-radius: 6px; border: 1px solid var(--rule); background: var(--panel); color: var(--ink); }
|
| 114 |
article h1 { font-size: 1.65rem; }
|
| 115 |
}
|
| 116 |
+
.masthead .links { margin: 6px 0 0; font: 500 12.5px/1.4 var(--f-mono); }
|
| 117 |
+
.masthead .links a { color: inherit; text-decoration: underline; text-underline-offset: 2px; opacity: .85; }
|
| 118 |
+
.masthead .links a:hover { opacity: 1; }
|
| 119 |
@media (prefers-reduced-motion: no-preference) { html { scroll-behavior: smooth; } }
|
| 120 |
</style>
|
| 121 |
|
|
|
|
| 123 |
<div class="masthead-inner">
|
| 124 |
<p class="brand"><span class="trits" aria-hidden="true"><i></i><i></i><i></i></span>bankml<span class="ver" id="ver">0.2.7</span></p>
|
| 125 |
<span class="tag">Verified low-bit inference for the CPU you already have — the documentation, as of <span id="asof">0.2.7</span>.</span>
|
| 126 |
+
<p class="links"><a href="https://github.com/cryptoAGI/bankml">GitHub</a> · <a href="https://huggingface.co/spaces/PYTHAI/bankml">Hugging Face</a> · <a href="https://deltaverse.pythai.net/bankml">why bankML</a> · <a href="#thesis">the thesis</a> · <a href="https://github.com/minaiml">minaiml</a> (<a href="https://huggingface.co/spaces/PYTHAI/minaiml">Space</a>)</p>
|
| 127 |
<div class="proof">gate record 0.3.6: the penalties — repeat, frequency and presence, token-identical to llama-server on <b>56 / 56</b> answers on each of three models, live <b>85 / 85</b> · gate record 0.3.5: JSON schemas — llama-server’s grammar on each template (<b>173 / 173</b> schemas × 3 templates against llama.cpp’s own code) and its answers token-identical on all five native models (<b>28 / 28</b>, <b>11 / 11</b>, <b>56 / 56</b> × 3), the content rule matched on <b>30,063 / 30,063</b> texts per template · <code>bankml create</code>: mindX’s persona layer, made from mindXtrain’s merged output, token-identical end to end (<b>27 / 27</b>, two ways) · gate record 0.3.4: mindX’s own model natively — <code>mindx-gen39</code> and SmolLM2-135M-Instruct (the Llama graph in F16) and Bonsai-1.7B (tied embeddings) token-identical to llama-server b11192: whole model <b>800 / 800</b> and <b>840 / 840</b> rows bit-exact, F16 products <b>552,268 / 552,268</b> elements bit-exact against ggml, seeded sampling <b>40 / 40</b> on each, conversations <b>9 / 9</b>, JSON mode <b>23 / 23</b> · JSON mode — answers under llama-server’s own <code>json_object</code> grammar token-identical, greedy and seeded (<b>23 / 23</b> answers on the 1-bit model, <b>13 / 13</b> on the ternary), and llama.cpp’s grammar sampler matched on <b>1,645 / 1,645</b> whole-vocabulary masks · a C API, <code>libbankml</code>, identical to <code>serve --native</code> · Savante answered by bankML’s own forward pass — whole conversations identical to llama-server on <b>9 / 9</b> turns (text, token counts, prompt-cache reuse) · own forward pass <b>1,064 / 1,064</b> rows bit-exact for the 1-bit and the ternary model · seeded sampling identical on <b>40 / 40</b> continuations · <b>8,188,239,872</b> ternary weights and <b>762 / 762</b> dot products bit-exact against llama.cpp b11192’s compiled library · one ternary token <b>0.222 s vs 2.139 s</b></div>
|
| 128 |
</div>
|
| 129 |
</header>
|
docs/install.md
CHANGED
|
@@ -353,6 +353,7 @@ oracles all do this.
|
|
| 353 |
| `--ctx N` | `4096` | both | the context: `-c N` for the spawned llama-server, `n_ctx` for the native engine. A prompt of `N` tokens or more is refused (0.3.8: with llama-server's `exceed_context_size_error` 400 body on `/v1`); an Ollama `num_ctx` above it is refused |
|
| 354 |
| `--spec-ngram` | off | llama.cpp | `--spec-type ngram-simple` in the spawned engine: n-gram speculative decoding, exact at temperature 0, opt-in because it measured within noise |
|
| 355 |
| `--slot-dir DIR` | — | both | llama.cpp mode: `--slot-save-path DIR` for the spawned engine. Native mode (0.3.8): `POST /slots/0?action=save\|restore\|erase` with `{"filename"}`, as llama-server's slot API; the directory is created at start. Either way, a restart restores the system prompt instead of computing it again. Without it, the slot actions answer 501 |
|
|
|
|
| 356 |
| `--native` | off | — | native mode |
|
| 357 |
| `--registry [DIR]` | off; bare: `$BANKML_FORKS`, else `~/.local/share/bankml/forks` | native | serve every model pinned in `DIR` by name, one resident at a time, each verified again when it loads; answer `/api/create`, `/api/copy`, `/api/delete` for derived models. A following argument that starts with `--` is not taken as `DIR` |
|
| 358 |
| `--keep-alive DUR` | `5m` | native | how long an `/api/*` request that names no `keep_alive` keeps the model resident: `"5m"`, `"1h30m"`, seconds, `0` (unload after the answer), negative (for good). OpenAI and llama-server endpoints keep the model resident |
|
|
|
|
| 353 |
| `--ctx N` | `4096` | both | the context: `-c N` for the spawned llama-server, `n_ctx` for the native engine. A prompt of `N` tokens or more is refused (0.3.8: with llama-server's `exceed_context_size_error` 400 body on `/v1`); an Ollama `num_ctx` above it is refused |
|
| 354 |
| `--spec-ngram` | off | llama.cpp | `--spec-type ngram-simple` in the spawned engine: n-gram speculative decoding, exact at temperature 0, opt-in because it measured within noise |
|
| 355 |
| `--slot-dir DIR` | — | both | llama.cpp mode: `--slot-save-path DIR` for the spawned engine. Native mode (0.3.8): `POST /slots/0?action=save\|restore\|erase` with `{"filename"}`, as llama-server's slot API; the directory is created at start. Either way, a restart restores the system prompt instead of computing it again. Without it, the slot actions answer 501 |
|
| 356 |
+
| `--allow-origin ORIGIN` | — | both | 0.3.9: one web page (`https://host[:port]`) whose scripts may call this gateway from a browser — the bankML Space's page, for example (`https://pythai-bankml.static.hf.space`). bankML answers that origin's CORS preflight (with Chrome's `Access-Control-Allow-Private-Network`) and adds `Access-Control-Allow-Origin` to its answers, streamed ones included; any other origin gets no CORS headers and a 403 preflight. The loopback `Host` and JSON-POST rules are unchanged |
|
| 357 |
| `--native` | off | — | native mode |
|
| 358 |
| `--registry [DIR]` | off; bare: `$BANKML_FORKS`, else `~/.local/share/bankml/forks` | native | serve every model pinned in `DIR` by name, one resident at a time, each verified again when it loads; answer `/api/create`, `/api/copy`, `/api/delete` for derived models. A following argument that starts with `--` is not taken as `DIR` |
|
| 359 |
| `--keep-alive DUR` | `5m` | native | how long an `/api/*` request that names no `keep_alive` keeps the model resident: `"5m"`, `"1h30m"`, seconds, `0` (unload after the answer), negative (for good). OpenAI and llama-server endpoints keep the model resident |
|
docs/modules/gguf.md
CHANGED
|
@@ -3,7 +3,7 @@
|
|
| 3 |
## Summary
|
| 4 |
|
| 5 |
`gguf.rs` reads the header of a GGUF v3 file and decides whether bankml may play it. The decision is the
|
| 6 |
-
**guard**: `play`, `refuse` with reasons, or `need_more` header bytes. It is a port of minaiml's Python
|
| 7 |
`gguf_guard.py` (same verdicts, same reasons, same JSON keys). The guard never reads tensor data.
|
| 8 |
|
| 9 |
The module also holds what the rest of bankml needs to reach tensor data safely: type names, block layouts,
|
|
|
|
| 3 |
## Summary
|
| 4 |
|
| 5 |
`gguf.rs` reads the header of a GGUF v3 file and decides whether bankml may play it. The decision is the
|
| 6 |
+
**guard**: `play`, `refuse` with reasons, or `need_more` header bytes. It is a port of [minaiml](https://github.com/minaiml)'s Python
|
| 7 |
`gguf_guard.py` (same verdicts, same reasons, same JSON keys). The guard never reads tensor data.
|
| 8 |
|
| 9 |
The module also holds what the rest of bankml needs to reach tensor data safely: type names, block layouts,
|
docs/thesis.md
CHANGED
|
@@ -132,7 +132,76 @@ bankML's discipline is to establish exactly that, against an independent referen
|
|
| 132 |
integrity, not proof: they bind the answer's text to a pinned file and to the request, between a client and its own
|
| 133 |
gateway ([research.md §3](research.md#3-verifiable-and-attested-inference)).
|
| 134 |
|
| 135 |
-
### II.5
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 136 |
|
| 137 |
It **inherits** ggml's formats, its kernels' semantics and llama-server's protocol. It **extends** them with an x86
|
| 138 |
ternary kernel the reference lacks, a gate in front of every answer, and receipts. It **breaks** with the convention
|
|
@@ -164,7 +233,8 @@ already shown identical.
|
|
| 164 |
result and not shipped ([`testing/experiments/`](../testing/experiments/)).
|
| 165 |
2. **The compiled reference is the specification, not its source.** Where the shipped library and its C source
|
| 166 |
disagree, the library wins (§II.3).
|
| 167 |
-
3. **A model that cannot be verified does not answer.** The guard ([`bankML/gguf.rs`](../bankML/gguf.rs)
|
|
|
|
| 168 |
from the header alone and with a reason, the three low-bit traps, including a file mainline loads and answers in
|
| 169 |
fluent nonsense. The pin ([`bankML/sha256.rs`](../bankML/sha256.rs)) refuses a file whose hash differs from its
|
| 170 |
provenance record.
|
|
@@ -313,6 +383,8 @@ Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Cameron, P., Jaggi, M., Alis
|
|
| 313 |
|
| 314 |
Cankaya, E. (2026). "Bit-Exact AI Inference Verification Without Performance Tradeoffs." [arXiv:2606.00279](https://arxiv.org/abs/2606.00279).
|
| 315 |
|
|
|
|
|
|
|
| 316 |
Codephreak, Professor and Magnusson, G. L. (2026). Design directives for bankML and the mindX runtime, recorded in the
|
| 317 |
project (project record, 2026-07-04 to 2026-09-28); quoted in [TECHNICAL.md](TECHNICAL.md#thesis--professor-codephreak-and-gregory-l-magnusson).
|
| 318 |
|
|
@@ -335,6 +407,9 @@ Surveys* 23(1): 5–48. [doi:10.1145/103162.103163](https://doi.org/10.1145/1031
|
|
| 335 |
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R. and Bengio, Y. (2016). "Binarized Neural Networks." *Advances in
|
| 336 |
Neural Information Processing Systems 29*. [Proceedings](https://papers.nips.cc/paper_files/paper/2016/hash/d8330f857a17c53d217014ee776bfd50-Abstract.html); preprint [arXiv:1602.02830](https://arxiv.org/abs/1602.02830).
|
| 337 |
|
|
|
|
|
|
|
|
|
|
| 338 |
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H. and Stoica, I. (2023).
|
| 339 |
"Efficient Memory Management for Large Language Model Serving with PagedAttention." *Proceedings of the 29th Symposium
|
| 340 |
on Operating Systems Principles (SOSP 2023)*. [arXiv:2309.06180](https://arxiv.org/abs/2309.06180).
|
|
@@ -344,9 +419,21 @@ Software* 39(2). [doi:10.1109/MS.2021.3073045](https://doi.org/10.1109/MS.2021.3
|
|
| 344 |
|
| 345 |
Li, F., Zhang, B. and Liu, B. (2016). "Ternary Weight Networks." [arXiv:1605.04711](https://arxiv.org/abs/1605.04711).
|
| 346 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 347 |
Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J. and Wei, F. (2024). "The Era of
|
| 348 |
1-bit LLMs: All Large Language Models are in 1.58 Bits." [arXiv:2402.17764](https://arxiv.org/abs/2402.17764).
|
| 349 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 350 |
Qwen Team (2025). "Qwen3 Technical Report." [arXiv:2505.09388](https://arxiv.org/abs/2505.09388).
|
| 351 |
|
| 352 |
Rastegari, M., Ordonez, V., Redmon, J. and Farhadi, A. (2016). "XNOR-Net: ImageNet Classification Using Binary
|
|
@@ -360,6 +447,14 @@ Thompson, K. (1984). "Reflections on Trusting Trust." *Communications of the ACM
|
|
| 360 |
Tseng, A., Chee, J., Sun, Q., Kuleshov, V. and De Sa, C. (2024). "QuIP#: Even Better LLM Quantization with Hadamard
|
| 361 |
Incoherence and Lattice Codebooks." *International Conference on Machine Learning (ICML 2024)*. [arXiv:2402.04396](https://arxiv.org/abs/2402.04396).
|
| 362 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 363 |
Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y. and Wei, F. (2023). "BitNet:
|
| 364 |
Scaling 1-bit Transformers for Large Language Models." [arXiv:2310.11453](https://arxiv.org/abs/2310.11453).
|
| 365 |
|
|
@@ -372,9 +467,18 @@ Wang, J., Zhou, H., Song, T. et al. (2025). "Bitnet.cpp: Efficient Edge Inferenc
|
|
| 372 |
Wei, J. et al. (2025). "T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge." *Proceedings of
|
| 373 |
EuroSys 2025*. [arXiv:2407.00088](https://arxiv.org/abs/2407.00088).
|
| 374 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 375 |
Zhu, C., Han, S., Mao, H. and Dally, W. J. (2017). "Trained Ternary Quantization." *International Conference on
|
| 376 |
Learning Representations (ICLR 2017)*. [arXiv:1612.01064](https://arxiv.org/abs/1612.01064).
|
| 377 |
|
|
|
|
|
|
|
| 378 |
*Notes on the references.* Author lists for the 2024–2026 preprints follow [research.md](research.md), which records
|
| 379 |
which details were re-fetched and which are as commonly cited. QuaRot and QuIP# are cited for the technique of
|
| 380 |
rotating activations by a Hadamard transform before quantization; the attribution of llama.cpp's own rotation to
|
|
|
|
| 132 |
integrity, not proof: they bind the answer's text to a pinned file and to the request, between a client and its own
|
| 133 |
gateway ([research.md §3](research.md#3-verifiable-and-attested-inference)).
|
| 134 |
|
| 135 |
+
### II.5 The contemporary field (survey of 6 October 2026)
|
| 136 |
+
|
| 137 |
+
This section maps the work bankML sits among, as found on 6 October 2026: every paper below was checked against its
|
| 138 |
+
arXiv abstract, and every repository's existence and last activity against GitHub on that date. Projects' own speed
|
| 139 |
+
figures are theirs, not reproduced here. A wider survey with the verifiable-inference field is in
|
| 140 |
+
[research.md](research.md).
|
| 141 |
+
|
| 142 |
+
#### II.5.1 Ternary and 1-bit models
|
| 143 |
+
|
| 144 |
+
Two programmes now train language models through the ternary constraint rather than quantizing after training. The
|
| 145 |
+
BitNet line (Microsoft Research) moved from the 1.58-bit recipe (Ma et al. 2024) to an openly released 2-billion-
|
| 146 |
+
parameter ternary model trained on 4 trillion tokens (Ma et al. 2025), then to 4-bit activations beside the 1-bit
|
| 147 |
+
weights (Wang, Ma and Wei 2024; 2025, the latter by a Hadamard transformation of the activations), and to distilling
|
| 148 |
+
full-precision models into 1.58 bits (Wu et al. 2025). Spectra (Kaushal et al. 2024) trained a suite of 54 models from
|
| 149 |
+
99 M to 3.9 B parameters to compare ternary "TriLMs" with float and post-training-quantized peers at equal size, and
|
| 150 |
+
Spectra 1.1 (Vaidhya et al. 2025) scaled TriLMs to 1.2 T tokens with scaling laws and its own inference kernel.
|
| 151 |
+
Alongside these, TernaryLLM (Chen et al. 2024) and OneBit (Xu et al. 2024) push quantization-aware training of
|
| 152 |
+
existing models to ternary and 1-bit weights, PTQ1.61 (Zhao et al. 2025) reaches below two bits after training,
|
| 153 |
+
ParetoQ (Liu et al. 2025) maps scaling laws across extreme bit widths, and MatMul-free language modelling (Zhu et al.
|
| 154 |
+
2024) removes matrix multiplication by ternary weights altogether.
|
| 155 |
+
|
| 156 |
+
The Bonsai models bankML serves come from PrismML and are dense Qwen3 networks trained to one bit (`Q1_0`) and to
|
| 157 |
+
ternary values (`Q2_0`, group 64, 2.25 bits per weight). PrismML's own format table
|
| 158 |
+
([docs.prismml.com](https://docs.prismml.com/download/formats)) records that group-64 `Q2_0` is in mainline llama.cpp,
|
| 159 |
+
that a group-128 legacy `Q2_0` is deprecated, and that its newer Ternary Bonsai 2 formats (`PQ2_0`, `PTQ1_0`) use a
|
| 160 |
+
rotated weight basis that needs an activation-side Walsh–Hadamard transform at run time — the file a stock engine
|
| 161 |
+
loads and answers in fluent nonsense, which is why bankML's guard refuses it (§III.2), and the same family of
|
| 162 |
+
rotation bankML reproduced for llama.cpp's quantized cache (§III.4).
|
| 163 |
+
|
| 164 |
+
#### II.5.2 Kernels and engines for ternary and 1-bit weights
|
| 165 |
+
|
| 166 |
+
| work | what it is | relation to bankML |
|
| 167 |
+
|---|---|---|
|
| 168 |
+
| [llama.cpp](https://github.com/ggml-org/llama.cpp) upstream | `Q1_0` ([#21273](https://github.com/ggml-org/llama.cpp/pull/21273); x86 AVX2+FMA in [#21636](https://github.com/ggml-org/llama.cpp/pull/21636)); `Q2_0` ([#24448](https://github.com/ggml-org/llama.cpp/pull/24448), NEON and scalar); an x86 `Q2_0` kernel needing AVX-VNNI ([#26348](https://github.com/ggml-org/llama.cpp/pull/26348), open); the older `TQ1_0`/`TQ2_0` ([#10010](https://github.com/ggml-org/llama.cpp/pull/10010)) | the reference: bankML reproduces b11192's compiled code bit for bit and adds the plain-AVX2 `Q2_0` kernel it lacks |
|
| 169 |
+
| [PrismML's fork](https://github.com/PrismML-Eng/llama.cpp) | kernels for the Bonsai 2 formats, e.g. `PQ2_0` AVX2/AVX-VNNI ([#206](https://github.com/PrismML-Eng/llama.cpp/pull/206)) | different formats; refused by bankML's guard until a kernel and an oracle exist |
|
| 170 |
+
| [bitnet.cpp](https://github.com/microsoft/BitNet) (Wang, Zhou, Song et al. 2024; 2025) | lookup-table and I2_S kernels for BitNet b1.58 on CPU; reports 2.37–6.17× on x86 | a different weight format (BitNet's), and a different exactness claim (lossless to its own model, not to an external reference) |
|
| 171 |
+
| [T-MAC](https://github.com/microsoft/T-MAC) (Wei et al. 2025) | lookup-table mixed-precision GEMM on CPU and NPU | the table-lookup alternative to bankML's `maddubs` arithmetic |
|
| 172 |
+
| Vec-LUT (Li et al. 2025) | vector table lookup for parallel ultra-low-bit inference on edge devices | the same direction as T-MAC, parallelised |
|
| 173 |
+
| Spectra 1.1's TriRun (Vaidhya et al. 2025) | a GPU kernel for packed ternary weights | GPU, not CPU |
|
| 174 |
+
|
| 175 |
+
#### II.5.3 Inference engines written in Rust
|
| 176 |
+
|
| 177 |
+
| engine | what it is | state on 2026-10-06 | exactness claim |
|
| 178 |
+
|---|---|---|---|
|
| 179 |
+
| [candle](https://github.com/huggingface/candle) (Hugging Face) | a minimalist ML framework; GGUF K-quants on CPU (AVX2/NEON), CUDA, Metal, WASM | active, the base most Rust LLM projects build on | none against llama.cpp found |
|
| 180 |
+
| [mistral.rs](https://github.com/EricLBuehler/mistral.rs) | an LLM server on candle: many architectures, ISQ, GPTQ/AWQ/HQQ/FP8 | active | none found |
|
| 181 |
+
| [burn](https://github.com/tracel-ai/burn) | a general tensor and deep-learning framework | active | not a GGUF engine |
|
| 182 |
+
| [Crane](https://github.com/lucasjinreal/Crane), [kalosm](https://github.com/floneum/kalosm), [cake](https://github.com/evilsocket/cake) | LLM/VLM engines and libraries on candle; cake distributes inference across devices | active | none found |
|
| 183 |
+
| [OxiLLaMa](https://github.com/cool-japan/oxillama) | a pure-Rust GGUF engine with its own AVX2/AVX-512/NEON kernels, including `TQ1_0`/`TQ2_0` and `Q1_0_G128` | active (alpha) | top-1 logit parity within a tolerance |
|
| 184 |
+
| [Frink](https://github.com/antonellof/frink) | a pure-Rust GGUF engine with quantized CPU, Metal and CUDA kernels and MoE | active | quantizer bytes identical on two formats |
|
| 185 |
+
| [Cera](https://github.com/hyeons-lab/cera) | a Rust-native GGUF engine (AVX2/AVX-512, NEON dotprod/i8mm, optional wgpu) | crate 0.6.3, 2026-09-25 | none stated |
|
| 186 |
+
| [llama-gguf](https://github.com/Lexmata/llama-gguf), [lm.rs](https://github.com/samuel-vitorino/lm.rs) | small engines, correctness-first / minimal | llama-gguf 2026-04; lm.rs inactive since 2024-10 | none stated |
|
| 187 |
+
| [bitnet-rs](https://github.com/lilyco-42/bitnet-rs), [bitnet-toy](https://github.com/tidynest/bitnet-toy) | BitNet b1.58 in Rust (a port of bitnet.cpp; a from-scratch teaching engine) | 2026-09 | unit tests against its own C++ baseline |
|
| 188 |
+
| [alice-aegis](https://github.com/Aefinity-AI/alice-aegis) | a `no_std` UEFI ternary engine with frozen integer semantics and SHA-256 receipts chaining the logits | 2026-10 | bit-identical across its own ISAs, not against an external reference |
|
| 189 |
+
| [ratchet](https://github.com/huggingface/ratchet), [tract](https://github.com/sonos/tract), [rten](https://github.com/robertknight/rten) | browser/WebGPU and ONNX inference | active | not GGUF engines |
|
| 190 |
+
| [rustformers/llm](https://github.com/rustformers/llm), [llama-cpp-rs](https://github.com/utilityai/llama-cpp-rs) | the first, archived (2024); the second, bindings to llama.cpp's C++ | — | inherit llama.cpp's arithmetic by calling it |
|
| 191 |
+
|
| 192 |
+
#### II.5.4 Where bankML stands among them
|
| 193 |
+
|
| 194 |
+
Three things in this survey are bankML's alone. It is the only engine found that reproduces the reference's
|
| 195 |
+
*compiled* library bit for bit — every weight, every dot product, every token of whole conversations — rather than
|
| 196 |
+
within a tolerance or against itself; alice-aegis shares the receipt idea and the exact integer discipline but checks
|
| 197 |
+
against its own builds. It is the only Rust engine found with ggml's group-64 ternary `Q2_0`, and its plain-AVX2 kernel
|
| 198 |
+
fills the x86 gap that upstream's open VNNI kernel leaves on CPUs without VNNI. And it is one zero-dependency binary
|
| 199 |
+
whose every answer carries a receipt. It is behind the field in breadth: candle and mistral.rs run far more
|
| 200 |
+
architectures and formats and run on GPUs, OxiLLaMa and upstream llama.cpp have NEON kernels for these formats, and
|
| 201 |
+
the BitNet and T-MAC kernels serve a different ternary format at speeds bankML has not been compared with. The
|
| 202 |
+
research programmes above produce the models; bankML's contribution is to run the ones in ggml's formats exactly.
|
| 203 |
+
|
| 204 |
+
### II.6 What bankML inherits, extends and breaks
|
| 205 |
|
| 206 |
It **inherits** ggml's formats, its kernels' semantics and llama-server's protocol. It **extends** them with an x86
|
| 207 |
ternary kernel the reference lacks, a gate in front of every answer, and receipts. It **breaks** with the convention
|
|
|
|
| 233 |
result and not shipped ([`testing/experiments/`](../testing/experiments/)).
|
| 234 |
2. **The compiled reference is the specification, not its source.** Where the shipped library and its C source
|
| 235 |
disagree, the library wins (§II.3).
|
| 236 |
+
3. **A model that cannot be verified does not answer.** The guard ([`bankML/gguf.rs`](../bankML/gguf.rs), a port of
|
| 237 |
+
the GGUF guard of [minaiml](https://github.com/minaiml), the authors' delivery layer for models on laptops and phones) refuses,
|
| 238 |
from the header alone and with a reason, the three low-bit traps, including a file mainline loads and answers in
|
| 239 |
fluent nonsense. The pin ([`bankML/sha256.rs`](../bankML/sha256.rs)) refuses a file whose hash differs from its
|
| 240 |
provenance record.
|
|
|
|
| 383 |
|
| 384 |
Cankaya, E. (2026). "Bit-Exact AI Inference Verification Without Performance Tradeoffs." [arXiv:2606.00279](https://arxiv.org/abs/2606.00279).
|
| 385 |
|
| 386 |
+
Chen, T., Li, Z., Xu, W. et al. (2024). "TernaryLLM: Ternarized Large Language Model." [arXiv:2406.07177](https://arxiv.org/abs/2406.07177).
|
| 387 |
+
|
| 388 |
Codephreak, Professor and Magnusson, G. L. (2026). Design directives for bankML and the mindX runtime, recorded in the
|
| 389 |
project (project record, 2026-07-04 to 2026-09-28); quoted in [TECHNICAL.md](TECHNICAL.md#thesis--professor-codephreak-and-gregory-l-magnusson).
|
| 390 |
|
|
|
|
| 407 |
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R. and Bengio, Y. (2016). "Binarized Neural Networks." *Advances in
|
| 408 |
Neural Information Processing Systems 29*. [Proceedings](https://papers.nips.cc/paper_files/paper/2016/hash/d8330f857a17c53d217014ee776bfd50-Abstract.html); preprint [arXiv:1602.02830](https://arxiv.org/abs/1602.02830).
|
| 409 |
|
| 410 |
+
Kaushal, A., Vaidhya, T., Mondal, A. K. et al. (2024). "Spectra: Surprising Effectiveness of Pretraining Ternary
|
| 411 |
+
Language Models at Scale." [arXiv:2407.12327](https://arxiv.org/abs/2407.12327).
|
| 412 |
+
|
| 413 |
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H. and Stoica, I. (2023).
|
| 414 |
"Efficient Memory Management for Large Language Model Serving with PagedAttention." *Proceedings of the 29th Symposium
|
| 415 |
on Operating Systems Principles (SOSP 2023)*. [arXiv:2309.06180](https://arxiv.org/abs/2309.06180).
|
|
|
|
| 419 |
|
| 420 |
Li, F., Zhang, B. and Liu, B. (2016). "Ternary Weight Networks." [arXiv:1605.04711](https://arxiv.org/abs/1605.04711).
|
| 421 |
|
| 422 |
+
Li, X., Yin, C., Wang, W. et al. (2025). "Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge
|
| 423 |
+
Devices." [arXiv:2512.06443](https://arxiv.org/abs/2512.06443).
|
| 424 |
+
|
| 425 |
+
Liu, Z., Zhao, C., Huang, H. et al. (2025). "ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization."
|
| 426 |
+
[arXiv:2502.02631](https://arxiv.org/abs/2502.02631).
|
| 427 |
+
|
| 428 |
+
Ma, S., Wang, H., Huang, S. et al. (2025). "BitNet b1.58 2B4T Technical Report." [arXiv:2504.12285](https://arxiv.org/abs/2504.12285).
|
| 429 |
+
|
| 430 |
Ma, S., Wang, H., Ma, L., Wang, L., Wang, W., Huang, S., Dong, L., Wang, R., Xue, J. and Wei, F. (2024). "The Era of
|
| 431 |
1-bit LLMs: All Large Language Models are in 1.58 Bits." [arXiv:2402.17764](https://arxiv.org/abs/2402.17764).
|
| 432 |
|
| 433 |
+
minaiml (software). "min ai ml — language models in miniature": delivery software for models on laptops and
|
| 434 |
+
phones, the origin of bankML's GGUF guard. [github.com/minaiml](https://github.com/minaiml) · [Hugging Face Space
|
| 435 |
+
PYTHAI/minaiml](https://huggingface.co/spaces/PYTHAI/minaiml).
|
| 436 |
+
|
| 437 |
Qwen Team (2025). "Qwen3 Technical Report." [arXiv:2505.09388](https://arxiv.org/abs/2505.09388).
|
| 438 |
|
| 439 |
Rastegari, M., Ordonez, V., Redmon, J. and Farhadi, A. (2016). "XNOR-Net: ImageNet Classification Using Binary
|
|
|
|
| 447 |
Tseng, A., Chee, J., Sun, Q., Kuleshov, V. and De Sa, C. (2024). "QuIP#: Even Better LLM Quantization with Hadamard
|
| 448 |
Incoherence and Lattice Codebooks." *International Conference on Machine Learning (ICML 2024)*. [arXiv:2402.04396](https://arxiv.org/abs/2402.04396).
|
| 449 |
|
| 450 |
+
Vaidhya, T., Kaushal, A., Jain, V. et al. (2025). "Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language
|
| 451 |
+
Models." [arXiv:2506.23025](https://arxiv.org/abs/2506.23025).
|
| 452 |
+
|
| 453 |
+
Wang, H., Ma, S. and Wei, F. (2024). "BitNet a4.8: 4-bit Activations for 1-bit LLMs." [arXiv:2411.04965](https://arxiv.org/abs/2411.04965).
|
| 454 |
+
|
| 455 |
+
Wang, H., Ma, S. and Wei, F. (2025). "BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs."
|
| 456 |
+
[arXiv:2504.18415](https://arxiv.org/abs/2504.18415).
|
| 457 |
+
|
| 458 |
Wang, H., Ma, S., Dong, L., Huang, S., Wang, H., Ma, L., Yang, F., Wang, R., Wu, Y. and Wei, F. (2023). "BitNet:
|
| 459 |
Scaling 1-bit Transformers for Large Language Models." [arXiv:2310.11453](https://arxiv.org/abs/2310.11453).
|
| 460 |
|
|
|
|
| 467 |
Wei, J. et al. (2025). "T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge." *Proceedings of
|
| 468 |
EuroSys 2025*. [arXiv:2407.00088](https://arxiv.org/abs/2407.00088).
|
| 469 |
|
| 470 |
+
Wu, X., Huang, S., Wang, W. et al. (2025). "BitNet Distillation." [arXiv:2510.13998](https://arxiv.org/abs/2510.13998).
|
| 471 |
+
|
| 472 |
+
Xu, Y., Han, X., Yang, Z. et al. (2024). "OneBit: Towards Extremely Low-bit Large Language Models." [arXiv:2402.11295](https://arxiv.org/abs/2402.11295).
|
| 473 |
+
|
| 474 |
+
Zhao, J., Zhang, M., Wang, M. et al. (2025). "PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training
|
| 475 |
+
Quantization Methods for Large Language Models." [arXiv:2502.13179](https://arxiv.org/abs/2502.13179).
|
| 476 |
+
|
| 477 |
Zhu, C., Han, S., Mao, H. and Dally, W. J. (2017). "Trained Ternary Quantization." *International Conference on
|
| 478 |
Learning Representations (ICLR 2017)*. [arXiv:1612.01064](https://arxiv.org/abs/1612.01064).
|
| 479 |
|
| 480 |
+
Zhu, R.-J., Zhang, Y., Abreu, S. et al. (2024). "Scalable MatMul-free Language Modeling." [arXiv:2406.02528](https://arxiv.org/abs/2406.02528).
|
| 481 |
+
|
| 482 |
*Notes on the references.* Author lists for the 2024–2026 preprints follow [research.md](research.md), which records
|
| 483 |
which details were re-fetched and which are as commonly cited. QuaRot and QuIP# are cited for the technique of
|
| 484 |
rotating activations by a Hadamard transform before quantization; the attribution of llama.cpp's own rotation to
|
docs/usage.md
CHANGED
|
@@ -247,6 +247,9 @@ In the next release (not yet tagged); each is checked against llama-server b1119
|
|
| 247 |
llama-server `-np 1`. Like llama-server, it keeps the states of other conversations in RAM (`BANKML_CACHE_RAM`, MiB;
|
| 248 |
0 off, -1 no limit; unset, 8192 MiB but at most a quarter of the memory available at load), so a conversation that comes back after
|
| 249 |
another does not recompute its whole history.
|
|
|
|
|
|
|
|
|
|
| 250 |
- **A smaller conversation memory (0.3.9, in progress).** `BANKML_CACHE_TYPE=q8_0` keeps the KV cache in q8_0 instead of f16,
|
| 251 |
about half the memory, so a long context fits on a small machine. It is llama.cpp's `--cache-type-k q8_0
|
| 252 |
--cache-type-v q8_0` exactly, with the Hadamard rotation llama.cpp applies around a quantized cache, and gives the
|
|
|
|
| 247 |
llama-server `-np 1`. Like llama-server, it keeps the states of other conversations in RAM (`BANKML_CACHE_RAM`, MiB;
|
| 248 |
0 off, -1 no limit; unset, 8192 MiB but at most a quarter of the memory available at load), so a conversation that comes back after
|
| 249 |
another does not recompute its whole history.
|
| 250 |
+
- **Your bankML from a web page (0.3.9).** `--allow-origin https://pythai-bankml.static.hf.space` lets the bankML
|
| 251 |
+
Space's page talk to your own `bankml serve` from your browser: your CPU, your verified model, a receipt on every
|
| 252 |
+
answer, nothing sent anywhere else. Only that one origin is answered with CORS headers.
|
| 253 |
- **A smaller conversation memory (0.3.9, in progress).** `BANKML_CACHE_TYPE=q8_0` keeps the KV cache in q8_0 instead of f16,
|
| 254 |
about half the memory, so a long context fits on a small machine. It is llama.cpp's `--cache-type-k q8_0
|
| 255 |
--cache-type-v q8_0` exactly, with the Hadamard rotation llama.cpp applies around a quantized cache, and gives the
|
index.html
CHANGED
|
@@ -125,6 +125,27 @@
|
|
| 125 |
.table td:first-child{color:#fff; font-weight:500; width:34%;}
|
| 126 |
.road li strong{color:rgb(var(--am));}
|
| 127 |
.hope p{font-size:18px; color:#eef;}
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 128 |
footer{max-width:780px; margin:0 auto; padding:10px 16px 90px; text-align:center; color:var(--quiet); font-size:14px;}
|
| 129 |
.social{display:flex; gap:22px; justify-content:center; align-items:center; margin:14px 0 10px;}
|
| 130 |
.social a{display:inline-flex; align-items:center; gap:8px; border:none; color:var(--muted);}
|
|
@@ -147,6 +168,43 @@
|
|
| 147 |
and 0.3.8 and 0.3.9 are being built on the way to the 0.4.0 milestone.</p>
|
| 148 |
</header>
|
| 149 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 150 |
<section>
|
| 151 |
<h2>bankML in one paragraph</h2>
|
| 152 |
<p>bankML runs language models on the computer you already have: a laptop, a small server, no graphics card
|
|
@@ -420,5 +478,6 @@
|
|
| 420 |
mount(layer);
|
| 421 |
})();
|
| 422 |
</script>
|
|
|
|
| 423 |
</body>
|
| 424 |
</html>
|
|
|
|
| 125 |
.table td:first-child{color:#fff; font-weight:500; width:34%;}
|
| 126 |
.road li strong{color:rgb(var(--am));}
|
| 127 |
.hope p{font-size:18px; color:#eef;}
|
| 128 |
+
#ask .mantra{font-size:19px;color:#fff;margin:.2em 0 .6em}
|
| 129 |
+
#ask .modes{display:grid;gap:6px;margin:8px 0}
|
| 130 |
+
#ask .modes label{color:var(--muted);cursor:pointer}
|
| 131 |
+
#ask .modenote{font-size:14px;color:var(--quiet);margin:.3em 0 .8em}
|
| 132 |
+
#ask .row{display:flex;flex-wrap:wrap;gap:8px;align-items:center;margin:8px 0}
|
| 133 |
+
#ask input[type=text],#ask #endpoint,#ask #model{flex:1 1 220px;min-width:0;background:rgba(0,0,0,.35);color:#fff;border:1px solid var(--rule);border-radius:8px;padding:8px 10px;font:14px ui-monospace,Menlo,monospace}
|
| 134 |
+
#ask textarea{flex:1 1 260px;min-width:0;background:rgba(0,0,0,.35);color:#fff;border:1px solid var(--rule);border-radius:8px;padding:9px 10px;font:15px system-ui,sans-serif;resize:vertical}
|
| 135 |
+
#ask button{background:rgba(var(--cy),.16);color:#fff;border:1px solid rgba(var(--cy),.5);border-radius:8px;padding:8px 14px;font:600 14px system-ui,sans-serif;cursor:pointer}
|
| 136 |
+
#ask button:hover{background:rgba(var(--cy),.3)}
|
| 137 |
+
#ask button:disabled{opacity:.5;cursor:wait}
|
| 138 |
+
#ask .status{font-size:13.5px;color:var(--quiet);flex:1 1 100%}
|
| 139 |
+
#ask details{flex:1 1 100%;font-size:14.5px;color:var(--muted)}
|
| 140 |
+
#ask details code{word-break:break-all}
|
| 141 |
+
#chat{display:grid;gap:8px;margin:12px 0 4px;max-height:420px;overflow-y:auto}
|
| 142 |
+
#chat .msg{white-space:pre-wrap;border-radius:10px;padding:9px 12px;line-height:1.5}
|
| 143 |
+
#chat .msg.user{background:rgba(var(--vi),.18);justify-self:end;max-width:85%;color:#fff}
|
| 144 |
+
#chat .msg.assistant{background:rgba(0,0,0,.35);border:1px solid var(--rule);color:#e8eef5}
|
| 145 |
+
#chat .meta{font:12.5px ui-monospace,Menlo,monospace;color:var(--quiet);margin:-4px 2px 4px}
|
| 146 |
+
#chat .meta.ok{color:#56d364}
|
| 147 |
+
#chat .meta.bad{color:#ff7b72}
|
| 148 |
+
#chat .meta.warn{color:rgb(var(--am))}
|
| 149 |
footer{max-width:780px; margin:0 auto; padding:10px 16px 90px; text-align:center; color:var(--quiet); font-size:14px;}
|
| 150 |
.social{display:flex; gap:22px; justify-content:center; align-items:center; margin:14px 0 10px;}
|
| 151 |
.social a{display:inline-flex; align-items:center; gap:8px; border:none; color:var(--muted);}
|
|
|
|
| 168 |
and 0.3.8 and 0.3.9 are being built on the way to the 0.4.0 milestone.</p>
|
| 169 |
</header>
|
| 170 |
|
| 171 |
+
|
| 172 |
+
<section id="ask">
|
| 173 |
+
<h2>Ask bankML</h2>
|
| 174 |
+
<p class="mantra" id="mantra"></p>
|
| 175 |
+
<div class="modes" role="radiogroup" aria-label="who answers">
|
| 176 |
+
<label><input type="radio" name="mode" value="local" checked> <b>Your own bankML</b> — verified, with a receipt, on your CPU (free)</label>
|
| 177 |
+
<label><input type="radio" name="mode" value="hf"> <b>A Hugging Face provider</b> — not bankML, no receipt, your inference quota</label>
|
| 178 |
+
</div>
|
| 179 |
+
<p class="modenote" id="modenote"></p>
|
| 180 |
+
<div id="localrow" class="row">
|
| 181 |
+
<input id="endpoint" aria-label="your bankml serve address" spellcheck="false">
|
| 182 |
+
<button id="connect" type="button">connect</button>
|
| 183 |
+
<span id="localstatus" class="status"></span>
|
| 184 |
+
<details>
|
| 185 |
+
<summary>Start your own bankML for this page (Linux, AVX2; free)</summary>
|
| 186 |
+
<ol>
|
| 187 |
+
<li><code>git clone https://github.com/cryptoAGI/bankml && cd bankml && ./install.sh</code> — builds bankML and imports and verifies Bonsai-8B.</li>
|
| 188 |
+
<li><code>./install.sh stop</code> (the installer's own serve does not allow web pages), then:<br>
|
| 189 |
+
<code>target/release/bankml serve .models/Bonsai-8B-Q1_0.gguf --fork ~/.local/share/bankml/forks/Bonsai-8B-Q1_0.gguf.FORK.json --native --allow-origin <span id="origin"></span></code></li>
|
| 190 |
+
<li>Press <b>connect</b>. Your browser may ask to let this page reach your machine: that is the request for your own bankML on 127.0.0.1, and nothing else.</li>
|
| 191 |
+
</ol>
|
| 192 |
+
</details>
|
| 193 |
+
</div>
|
| 194 |
+
<div id="hfrow" class="row" hidden>
|
| 195 |
+
<input id="model" aria-label="provider model (owner/name)" spellcheck="false">
|
| 196 |
+
<button id="signin" type="button">Sign in with Hugging Face</button>
|
| 197 |
+
<button id="signout" type="button" hidden>sign out</button>
|
| 198 |
+
<span id="who" class="status"></span>
|
| 199 |
+
</div>
|
| 200 |
+
<div id="chat" aria-live="polite"></div>
|
| 201 |
+
<div class="row ask">
|
| 202 |
+
<textarea id="message" rows="2" placeholder="Ask bankML about itself, its exactness, its speed… (Ctrl+Enter)"></textarea>
|
| 203 |
+
<button id="send" type="button">Ask</button>
|
| 204 |
+
</div>
|
| 205 |
+
<p id="askstatus" class="status"></p>
|
| 206 |
+
</section>
|
| 207 |
+
|
| 208 |
<section>
|
| 209 |
<h2>bankML in one paragraph</h2>
|
| 210 |
<p>bankML runs language models on the computer you already have: a laptop, a small server, no graphics card
|
|
|
|
| 478 |
mount(layer);
|
| 479 |
})();
|
| 480 |
</script>
|
| 481 |
+
<script type="module" src="bankml-chat.js"></script>
|
| 482 |
</body>
|
| 483 |
</html>
|
sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg and b/sAGI/voice/cache/savante/01816c6078fe013e4876519f.ogg differ
|
|
|
sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg and b/sAGI/voice/cache/savante/018f79aceb48108e67b45102.ogg differ
|
|
|
sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg and b/sAGI/voice/cache/savante/02aaba262b4dcf8dcace911f.ogg differ
|
|
|
sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg and b/sAGI/voice/cache/savante/02fee06d31832ce2c65ad09f.ogg differ
|
|
|
sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg and b/sAGI/voice/cache/savante/0307f8958bd59e22deb29ca2.ogg differ
|
|
|
sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg and b/sAGI/voice/cache/savante/03181886c7327d6c927ff7d2.ogg differ
|
|
|
sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg and b/sAGI/voice/cache/savante/048c96f0872ba00331469160.ogg differ
|
|
|
sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg and b/sAGI/voice/cache/savante/0523f2db73f1278da1ed0352.ogg differ
|
|
|
sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg and b/sAGI/voice/cache/savante/052bcb9d71abb283bdf471dc.ogg differ
|
|
|
sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg and b/sAGI/voice/cache/savante/064a6ebaf4d8116c27c04b82.ogg differ
|
|
|
sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg and b/sAGI/voice/cache/savante/06964651d6f65fd28460040c.ogg differ
|
|
|
sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg and b/sAGI/voice/cache/savante/0817b60228e9e594923e8c31.ogg differ
|
|
|
sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg and b/sAGI/voice/cache/savante/0a1a6bcf8464335339882f36.ogg differ
|
|
|
sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg and b/sAGI/voice/cache/savante/0a5088e62c596d725b72cc56.ogg differ
|
|
|
sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg and b/sAGI/voice/cache/savante/0a59208964fc50565f0e2272.ogg differ
|
|
|
sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg and b/sAGI/voice/cache/savante/0c4e43f59479a6a35e6cd8f2.ogg differ
|
|
|
sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg and b/sAGI/voice/cache/savante/0c8fa0d6541451d4c493c629.ogg differ
|
|
|
sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg and b/sAGI/voice/cache/savante/0cc164083b4fabab03e781ba.ogg differ
|
|
|
sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg and b/sAGI/voice/cache/savante/0d02408388431bed4f42574b.ogg differ
|
|
|
sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg and b/sAGI/voice/cache/savante/0d968446ff0cb0e8aa6d468d.ogg differ
|
|
|
sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg and b/sAGI/voice/cache/savante/0f861912836521a0d0b936d2.ogg differ
|
|
|
sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg and b/sAGI/voice/cache/savante/10343befd6ff6c093c705bcd.ogg differ
|
|
|
sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg and b/sAGI/voice/cache/savante/111b2a2efe26a34ce04ebda1.ogg differ
|
|
|
sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg and b/sAGI/voice/cache/savante/11e59c53876c901b98f936af.ogg differ
|
|
|
sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg and b/sAGI/voice/cache/savante/125e136fdaa7e82376a7e520.ogg differ
|
|
|
sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg and b/sAGI/voice/cache/savante/12c460cad05f9593ee368047.ogg differ
|
|
|
sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg and b/sAGI/voice/cache/savante/12d890a48faa44a9daed02e7.ogg differ
|
|
|
sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg and b/sAGI/voice/cache/savante/131c4c56f6d6867d5bee90b9.ogg differ
|
|
|
sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg and b/sAGI/voice/cache/savante/13a1afdd97f297efa6d8e942.ogg differ
|
|
|
sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg and b/sAGI/voice/cache/savante/1401fd80589a41cf378f4d01.ogg differ
|
|
|
sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg and b/sAGI/voice/cache/savante/155e3cc0b8f953e1fc60a503.ogg differ
|
|
|
sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg and b/sAGI/voice/cache/savante/1602ef1c516b4ec287c6e18f.ogg differ
|
|
|
sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg and b/sAGI/voice/cache/savante/1949c558865d80dc2c7f77a7.ogg differ
|
|
|
sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg and b/sAGI/voice/cache/savante/19acaa1e7dc06a7fcc88c3df.ogg differ
|
|
|
sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg and b/sAGI/voice/cache/savante/19cf9ce143b927fd946ea744.ogg differ
|
|
|
sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg
CHANGED
|
Binary files a/sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg and b/sAGI/voice/cache/savante/1ad61988f3c2b9cef1b85a45.ogg differ
|
|
|