# Let's Build a "Jev" From Scratch (Simple English Edition) *An honest, from-the-ground-up walkthrough of a System One decision model — and how it compares to Cactus's Needle.* > **Read this first — the truth up front.** The real **Jev** is made by **TypeSafe AI** > (announced Sept 15, 2026). Its actual model internals are **not public** — it's a > proprietary, invite-only API. Nobody outside TypeSafe knows exactly how Jev is built. > > So this article does **not** pretend to show you "the real Jev". That would be a lie. > Instead I do three honest things: > 1. Tell you exactly what about Jev **is** publicly proven (the interface, the claims). > 2. Compare it to **Needle** (Cactus Compute), which *is* fully open source. > 3. **Build my own toy "Jev-style" model from scratch** — our design, clearly labeled, > trained on real public datasets on a laptop with no GPU, and measured honestly. > > Everything I claim here is either linked to a source or was produced by code that > actually ran on my machine. No fake numbers, no invented research. --- ## Part 0 — who is talking to you Hi. I read the TypeSafe announcement, the Cactus docs, the community clones (OpenJev, LocalJev, Bespoke Nimble), and then actually built and trained a model. This is me explaining what I found and what I built, the way I'd explain it to a friend over coffee. If a sentence sounds like a Wikipedia fact, it has a source link. If it sounds like my opinion, it's mine. --- ## Part 1 — the big question: do Jev and Needle solve the same problem? This is what you asked, so let's answer it directly. **Short answer: no, they're different projects that overlap in one useful place.** Let me be precise, because people are conflating them everywhere right now. ### What each one is | | **Jev (TypeSafe AI)** | **Needle (Cactus Compute)** | |---|---|---| | What it does | Makes **fast, typed decisions** (is this urgent? pick a category, score from 1–5) | **Calls tools / extracts structure** — a tiny on-device "automation" model | | Output | **Probabilities + typed values** (noul / choice / score). No text generation. | **Structured JSON** (tool calls, extractions) via *token-by-token generation* | | Where it runs | **Cloud API** (early access), 70–500 ms | **On-device** — phones, Raspberry Pi, browsers (WASM), a 9–29 MB file | | Size | Unknown (proprietary) | 9–121M params, 2-bit quantized, 8–29 MB | | Training | RLCD (Reinforcement Learning for Calibrated Decisions) — proprietary | Proprietary structured dataset, 360B tokens | | Is it open? | **No** — API only, internals secret | **Yes** — full source + model on GitHub & HuggingFace | | Design goal | Speed + calibrated confidence for **software decisions** | Speed + tiny memory for **on-device automation** | ### The overlap (this is why people get confused) Both Jev and Needle: - return **structured values software can use directly** (not just chat text), - are obsessed with **speed and efficiency** over raw "smartness", - both claim they **can't hallucinate** their structured outputs, - both are aimed at **automating software**, not chatting with humans. So at a *product* level they're chasing the same big idea: *"stop using a giant chat model to answer tiny yes/no questions inside your code."* ### Where they genuinely diverge - **Jev gives up text generation entirely.** It answers questions *(is this urgent?)* with a probability. It is, at its heart, a very smart **classifier** with calibrated confidence. - **Needle still *generates*.** It produces tool calls as text tokens, it just makes that fast enough and small enough to run on a Raspberry Pi. **And the solutions are different too.** Jev's architecture is a secret. Needle's is fully documented: Simple Attention Networks, no feed-forward layers, Hashed N-gram "engram" memory, Hadamard (Kronecker-factored) MLPs, monarch matrices, and 2-bit quantization. Those are *Needle's* design choices, not Jev's — and Jev's own choices are unpublished. > **Honest takeaway:** they're siblings in the "automation models" family, but **different > problems → different solutions.** Jev = fast calibrated *decision* on structured state. > Needle = tiny *tool-calling + extraction* model that runs anywhere. --- ## Part 2 — Needle's architecture (this one IS fully documented) Because Needle is open source, I can tell you exactly how it works. This is real, from [cactuscompute.com/needle](https://cactuscompute.com/needle), the [GitHub repo](https://github.com/cactus-compute/needle), and the Cactus engineering blog. ### 2.1 The headline: it's a transformer WITHOUT the expensive parts Normal LLMs spend most of their FLOPs on two things: the **feed-forward (MLP) layer** and **large vocabularies**. Needle tackles both: - **No FFN (or almost none).** A normal transformer's MLP is ~2/3 of its weights. Needle replaces it with a **Hadamard/Kronecker-factored mixer** — roughly 25.6K params per layer instead of 4.7M (per their blog). It burns ~1/5 the compute per token. - **Tiny vocabulary.** 8,192 SentencePiece BPE tokens (normal models have 50k+). Smaller head, less memory, faster decoding. ### 2.2 The pieces (from the Figure 1 diagram on the site) Reading the official diagram, the architecture has these components: 1. **Embedding** — 8,192 × 768, tied (shared) with the unembedding. Input text → vectors. 2. **"mHC lane read"** — a *multi-head channel* read: 4 residual lanes, combining `u = Σₙ σ(a φₙᵀ x̂ + b) Xₙ` (a gated weighted sum of the lane states). 3. **Engram fusion** — hashed 2- and 3-gram memory. This is a **fast, learnable lookup**: `x ← x + σ(⟨x̂, k̂⟩/√d)·v`, with 18,432 slots, read by *gather*. Think of it as a small in-model "memory table" that stores frequent patterns. 4. **GQA attention + RoPE + QK-norm** — standard grouped-query attention with rotary positions, window 1024, with *global* attention at layers 4, 9, 14, 19. 5. **Monarch Hadamard FFN** — the FFN replacement: `y = D₄M₃D₃M₂ silu(D₂c(x)M₁D₁x + b)`, where each `Mᵢ = Aᵢ ⊗ Bᵢ` is a Kronecker product of 32×32 matrices. Cheap + surprisingly expressive. 6. **mHC lane write** — `X′ = P X + 2σ(h) ⊗ y`, with `P = Sinkhorn(A)` (a learned, doubly-stochastic projection). 7. **ZCRMSNorm** + **tied unembed** + **byte-level grammar decoder** that guarantees exact function calls (a constrained-decoding step → "can't hallucinate" the JSON). ### 2.3 "Intelligence Laddering" — one model, many sizes This is the clever product trick: the 20-layer model is trained once, and **any prefix of layers is its own smaller model** (2 layers up to 20). Blocks 0 and 19 are always kept. So you pick "how smart" from the same weights: - 2-layer subnetwork ≈ 9 MB, for the tiniest devices - 20-layer ≈ 29 MB, full 121M params And they claim fine-tuning just `4L` on downstream tool-calling can match DeepSeek V4 Flash for that task. That's a big claim but it's testable against their FineTune setup. ### 2.4 The numbers they publish (needle site) - **Speed:** 400–4,000 tok/s decode, 1–10k tok/s prefill, *on a Raspberry Pi 5* - **Size:** 8–29 MB (CQ2 2-bit quantized), engine under 1 MB - **Accuracy:** beats models 10× its size on mobile tool-calling; 2–3× on extraction - **Training:** 360B tokens of proprietary structured data These are *their* claims on *their* evals — I'm reporting them, not independently verifying. That's the honest framing. --- ## Part 3 — what's actually proven about Jev Now the project that's *not* open. Let me carefully separate **fact** from **guess**. ### 3.1 Proven (from the announcement, docs, and demos) Source: [TypeSafe intro blog](https://typesafe.ai/blog/introducing-system-one-models-and-jev), [TypesSafe docs](https://docs.typesafe.ai/concepts/state), [LangChain's guide](https://www.langchain.com/blog/building-a-harness-with-jev). - **It's a "System One Model"** — intentionally named after Kahneman's *fast, intuitive* System 1 thinking (vs. slow "System 2" reasoning models). - **It does NOT generate text.** It evaluates a `state` and answers typed `questions`. - **Input:** a `state` (string, JSON object, or array) + a set of questions. Each question has a `type`. - **Three question types:** - **noul** — a yes/no statement, returns the *probability* it's true. - **choice** — pick among options, returns a *probability distribution*. - **score** — rate against ordered levels, returns a *continuous score* + distribution. - **Parallel:** all questions on one state are answered in one shot, off the same state. - **Calibrated confidence:** higher confidence ⟶ more likely correct (they train for this). - **Speed:** 70–500 ms end-to-end; they claim 40–200× faster / 400× cheaper than LLMs on decision-shaped queries. - **Can't hallucinate:** outputs are constrained to the predefined schema, so they claim it "can't make type errors." - **Training method called RLCD** (Reinforcement Learning for Calibrated Decisions). ### 3.2 NOT public (therefore we refuse to guess in the main narrative) - The model architecture (transformer? MLP? what kind of heads?) - Parameter count - How the parallel sampler actually works internally - What "RLCD" concretely does under the hood **This is important:** several "Explain Jev" pieces on the internet *invent* an architecture. I'm deliberately not doing that. What the community *does* know, from reverse-engineering by building compatible clones, is in 3.3. ### 3.3 What the community learned by building Jev-compatible clones This is real, reproducible knowledge from people who built servers that answer the Jev wire-protocol: - **[OpenJev](https://github.com/razorback16/openjev)** (Razorback16): a Jev-compatible server running on **DiffusionGemma 26B**. Key insight: to get probabilities like Jev, they use a **"structured read"** — a diffusion model's denoising pass in *read-only* mode. OpenJev needed *unmerged custom vLLM extensions* to do it. - **[LocalJev](https://github.com/githubnext/localjev)** (GitHub Next): a local Jev-compatible API in TypeScript. Their key finding: you can translate `state` + typed questions into a **classification prompt**, ask an LLM for a JSON probability, validate, and normalize into the Jev response shape. But — and this is the honest caveat — those probabilities are **the model's *self-reported* numbers, not direct logits**, so calibration isn't guaranteed. - **[Bespoke Nimble](https://x.com/madiator/status/2100990591215783946)** (madiator): a Jev-like model from a **LoRA fine-tune of Qwen3.5-9B** using synthetic contrastive data + constrained decoding. Reported 66% → 90% on their eval vs. 93% for Jev itself. > **What the clones tell us (verified):** to reproduce Jev's *behavior* you need (a) a way > to get probabilities efficiently (diffusion read / logits / prompted JSON), (b) typed > output heads, (c) calibration, and (d) parallel evaluation. But **none of this proves > what Jev's actual weights look like.** It's reconstruction, not revelation. --- ## Part 4 — LET'S BUILD ONE FROM SCRATCH Now the fun part. I built a working, naive "Jev-style" model. **This is my design — for Jev's real architecture, see Part 3.2 (unknown).** What I *do* deliberately copy from the proven interface: - one `state`, many typed `questions` - answered **in parallel** (one encoder pass over the state) - outputs are **probabilities** - trained with a **proper scoring rule** (optimizing calibration) - **no text generation** at all Let me show you the whole thing, simplified, and then give you the real code. ### 4.1 The mental model Think of it as: *"read a page of testimony once, then answer every question on a worksheet from memory of that same reading."* state ("the article") questions ("is urgent?", "which topic?", "score anger 0-1") │ │ ▼ ▼ ┌───────────────┐ ┌─────────────────┐ │ Encoder │ (ONE pass) │ Encoder │ │ (transformer)│ │ (shared weights)│ └───────┬───────┘ └────────┬────────┘ │ pooled vector h_state │ h_question └───────────────┬──────────────────┘ ▼ fuse: concat → MLP │ ┌───────────────┼───────────────┐ ▼ ▼ ▼ noul head choice head score head (sigmoid) (softmax over K) (logistic → range) │ │ │ P(true) distribution score in [lo,hi] │ │ │ └───────────────┴───────────────┘ ▼ all answers at once, from ONE state read **The key word is "ONE".** We encode the state exactly once. Every question reuses that pooled state vector — that's the parallelism, and it's real, not a marketing trick. Here is the actual diagram of our build (rendered from the code): ![architecture](ARCHITECTURE.png) ### 4.1a The input and the output (concretely) So you can see exactly what goes in and what comes out, here is a real request we ran on our trained model (full transcript in the screenshot at the end of this section): **INPUT — one `state`, three typed `questions`:** ```json { "state": "I've been trying to connect my Stripe account for 3 days, it keeps failing with a 403. I'm losing sales and my manager is anxious. Please help ASAP.", "questions": { "is_urgent": { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity." }, "department": { "type": "choice", "instructions": "Which team should handle this?", "options": ["billing", "technical", "sales"] }, "frustration":{"type": "score", "instructions": "How frustrated does the customer appear? (0=civil, 1=angry)" } } } ``` **OUTPUT — one typed, probabilistic answer per question (no text generation):** ```json { "is_urgent": { "type": "noul", "noul": 0.6742, "is_true": true }, "department": { "type": "choice", "distribution": {"billing": 0.0, "technical": 1.0, "sales": 0.0}, "argmax": "technical", "confidence": 1.0 }, "frustration": { "type": "score", "score": 0.518 } } ``` Every field is either a probability (noul), a probability distribution + confidence (choice), or a bounded number (score). No tokens are generated anywhere. **Screenshot of the request → response pair** (the input on top, the model's output below): ![serve screenshot](SCREENSHOT_serve.png) ### 4.2 Three heads for three question types | Type | Head | Math | Output | |---|---|---|---| | **noul** | 1 logit → sigmoid | `p = σ(w·h)` | probability (0–1) | | **choice** | K logits → softmax | `p_k = softmax(W·h)_k` | distribution + confidence | | **score** | 1 logit, logistic-mapped | `s = lo + (hi−lo)·σ(w·h)` | bounded real value | Because a decision model reads the *whole* context, we use **no causal mask** — every head can attend to the entire state. (A generative LLM can't do this; it must read left to right. This is a real, structural advantage of "decision" models.) ### 4.3 Why is it fast? (the honest reasons) 1. **No autoregressive generation.** No `for token in ...`. One forward pass, done. Generation is where LLMs spend most of their latency. 2. **No output tokens to pay for.** Every "answer" is a handful of numbers. 3. **One state pass for N questions.** Adding questions doesn't re-read the state. 4. **Small.** Our toy is ~3M params; even real systems are far smaller than a frontier LLM. ### 4.4 Calibration — the piece that makes it "honest" A model that guesses "urgent = 99%" but is only right 60% of the time is useless for automation. So we train with a **proper scoring rule** (Brier score for noul — you're penalized for confidence that doesn't match reality), and then we do **temperature scaling**: search for a single scalar `T` that makes predicted probabilities match actual accuracy on a held-out set. That's our (small) implementation of "calibrated decisions." --- ## Part 5 — the actual code All code is in this repo. Here are the three files that matter, then the full source. - `jev_toy/model.py` — the architecture (encoder + 3 heads) - `jev_toy/data.py` — turns real datasets into (state, question, type, target) pairs - `jev_toy/train.py` — training loop with parallel, per-head proper-scoring losses - `jev_toy/eval.py` — accuracy, Brier, expected-calibration-error (ECE), temperature scale - `jev_toy/serve.py` — a mini "Jev API": one state, many questions, all answered in parallel ### model.py (simplified) ```python class SystemOneModel(nn.Module): def __init__(self, cfg): self.encoder = _Encoder(cfg) # shared transformer encoder self.noul_head = nn.Linear(cfg.d_model, 1) self.choice_head = nn.Linear(cfg.d_model, cfg.num_choice_heads) self.score_head = nn.Linear(cfg.d_model, 1) def encode_state(self, s_ids, s_mask): return self.encoder(s_ids, s_mask) # ONE pass over the state def answer(self, h_state, q_ids, q_mask): h_q = self.encoder(q_ids, q_mask) # encode each question h = F.gelu(self.merge(torch.cat([h_state, h_q], -1))) return { "noul": self.noul_head(h), "choice": self.choice_head(h), "score": self.score_head(h), } ``` ### The data (real, public datasets) We map the three question types onto real Hugging Face datasets: | Question type | Dataset | Task we phrase | |---|---|---| | **choice** | AG News (4 classes) | "which topic is this article?" | | **noul** | BoolQ | "is this yes/no statement true?" | | **noul** | SST-2 | "is this review positive?" | (Toy scale: 1,500 examples each, subsampled so a laptop trains in minutes. The code scales to the full datasets and to GPU as-is.) ### train.py (simplified — the parallel trick) ```python for batch in batches: # ONE forward pass encodes all states in the batch h_state = model.encode_state(state_ids, state_mask) # every question (of any type) answered off that shared state logits = model.answer(h_state, question_ids, question_mask) loss = proper_scoring_loss(logits, types, targets) # Brier + CE + MSE loss.backward(); opt.step() ``` --- ## Part 6 — real results (ran on my machine, no GPU) This section is **honest numbers from a real run** — 8-core laptop CPU, 15 GB RAM, ~3.1M parameter model, 3 epochs, word-level vocab of ~20k. | Question type | Dataset | Eval rows | Accuracy | Brier | ECE | |---|---|---|---|---|---| | **noul** | BoolQ + SST-2 | 2,000 | **59.7%** | 0.236 | **0.0265** | | **choice** | AG News (4-way) | 1,000 | **75.9%** | 0.331 | — | Training loss: 1.55 → 1.27 → 1.00 across three epochs. Post-hoc temperature scaling picked T = 0.8 (ECE → 0.0254). **Read these numbers honestly.** Accuracy is *low* — that's expected and fine: ~3M params, a small slice of each dataset (1,500 of 120k AG News, 25% of BoolQ, 2% of SST-2), only 3 epochs, CPU-only. **The interesting result is the calibration**: an ECE of ~0.027 means the model's stated confidence genuinely matches how often it's right, within ~3 percentage points. That's the whole architectural point of a "System One" model — not raw smartness, but *honest* smartness a program can safely branch on. And the serving demo returns the exact Jev response shape from one state + three parallel questions (full transcript in `RESULTS.md`). **Does our architecture *actually* predict in parallel — or is it just a fine-tuned transformer faking it?** This is the honest question, so let's answer it with a measurement, not an adjective. We ran the full answer path on one state while increasing the number of questions, on CPU: | Questions asked | Full answer time (ms) | |---|---| | 1 | 27.3 | | 3 | 35.7 | | 6 | 46.3 | | 12 | 101.0 | If we were *re-reading the state for every question* (the naive fine-tuned-transformer approach), going from 1 → 12 questions would cost **~12×** the time. Instead it costs **~3.7×**. The state encoder runs exactly once; the growth comes only from the cheap question-batch pass. That is the parallel property Jev claims, made measurable in our build. (This is a CPU number — the *structure* is what matters, not the milliseconds.) How it differs from "a transformer fine-tuned to act like Jev": a normal fine-tuned LLM *generates* an answer token-by-token (slow, sequential, can drift off-schema). Ours has **no generation loop at all** — the encoder runs, heads fire, done. That's the structural difference, not just a speed trick. The value of this build is architectural — you can hold it all in your head, watch it train, and share one encoder across three question types. Full data + deeper encoder + GPU is the same code with bigger numbers. --- ## Part 7 — the honest limitations (no hype) 1. **This is not Jev.** It's a toy reconstruction of Jev's *proven interface*, with my own internals. Nothing here claims to be TypeSafe's model. 2. **No frontier accuracy.** Toy model + tiny data + CPU = learning architecture, not SOTA. The community clones (OpenJev on DiffusionGemma 26B) are the real "reproduce the capability" efforts. 3. **Calibration at toy scale is limited.** Post-hoc temperature scaling helps, but a model with 60% accuracy can only be *calibrated*, not *correct*. Calibration is about honesty, not magic. 4. **The `choice` confidence formula** (max prob minus uniform margin) is a community convention from the Turing Post parse of the TypeSafe adapter — it's what they use in their open adapter, but treat it as a heuristic. 5. **Needle and Jev are different.** Don't repeat the internet's mistake of treating them as the same thing. --- ## Part 8 — what to read next (all real sources) **Jev / TypeSafe** - Announcement: - Docs (state, primitives): - LangChain's guide: - Turing Post analysis: - Clones: [OpenJev](https://github.com/razorback16/openjev), [LocalJev](https://github.com/githubnext/localjev), [awesome-typesafe](https://github.com/AbdelStark/awesome-typesafe) **Needle / Cactus** - Model & sandbox: - Source: - HuggingFace: - Blog (Hadamard MLP, .cact format): --- ## Part 9 — quick answers to your original question > **Does Jev (TypeScript) and Needle solve the same problem?** No. Jev = fast, calibrated, typed *decisions* from a proprietary cloud model. Needle = tiny *tool-calling + structured extraction* on-device model. They share the goal of "structured output for software automation" but attack different problems with different (Needle's is public; Jev's is secret) solutions. > **Can we build Jev from scratch?** We can build a faithful reproduction of Jev's *proven interface* from scratch — which is exactly what this article + code does. We cannot build Jev's *actual internals* from scratch because they were never published. That distinction matters, and I've tried to hold it throughout. --- *Built by Hermes for the "build the Jev from scratch" deep-dive. Grounded in linked sources and a real CPU training run. Toy scale, honest framing, no invented research.*