Download ARTICLE.md from azharmo/build-jev-from-scratch: direct link, hf CLI and curl.
- Browser
- Download file 24.3 kB
-
https://huggingface.co/azharmo/build-jev-from-scratch/resolve/f48dabc6e187bfc9c7b5a9c2699c5320667730d4/ARTICLE.md
- Command line
-
hf download hf://azharmo/build-jev-from-scratch@f48dabc6e187bfc9c7b5a9c2699c5320667730d4/ARTICLE.md
-
curl -L -o ARTICLE.md https://huggingface.co/azharmo/build-jev-from-scratch/resolve/f48dabc6e187bfc9c7b5a9c2699c5320667730d4/ARTICLE.md
Let's Build a "Jev" From Scratch (Simple English Edition)
An honest, from-the-ground-up walkthrough of a System One decision model β and how it compares to Cactus's Needle.
Read this first β the truth up front. The real Jev is made by TypeSafe AI (announced Sept 15, 2026). Its actual model internals are not public β it's a proprietary, invite-only API. Nobody outside TypeSafe knows exactly how Jev is built.
So this article does not pretend to show you "the real Jev". That would be a lie. Instead I do three honest things:
- Tell you exactly what about Jev is publicly proven (the interface, the claims).
- Compare it to Needle (Cactus Compute), which is fully open source.
- Build my own toy "Jev-style" model from scratch β our design, clearly labeled, trained on real public datasets on a laptop with no GPU, and measured honestly.
Everything I claim here is either linked to a source or was produced by code that actually ran on my machine. No fake numbers, no invented research.
Part 0 β who is talking to you
Hi. I read the TypeSafe announcement, the Cactus docs, the community clones (OpenJev, LocalJev, Bespoke Nimble), and then actually built and trained a model. This is me explaining what I found and what I built, the way I'd explain it to a friend over coffee. If a sentence sounds like a Wikipedia fact, it has a source link. If it sounds like my opinion, it's mine.
Part 1 β the big question: do Jev and Needle solve the same problem?
This is what you asked, so let's answer it directly.
Short answer: no, they're different projects that overlap in one useful place.
Let me be precise, because people are conflating them everywhere right now.
What each one is
| Jev (TypeSafe AI) | Needle (Cactus Compute) | |
|---|---|---|
| What it does | Makes fast, typed decisions (is this urgent? pick a category, score from 1β5) | Calls tools / extracts structure β a tiny on-device "automation" model |
| Output | Probabilities + typed values (noul / choice / score). No text generation. | Structured JSON (tool calls, extractions) via token-by-token generation |
| Where it runs | Cloud API (early access), 70β500 ms | On-device β phones, Raspberry Pi, browsers (WASM), a 9β29 MB file |
| Size | Unknown (proprietary) | 9β121M params, 2-bit quantized, 8β29 MB |
| Training | RLCD (Reinforcement Learning for Calibrated Decisions) β proprietary | Proprietary structured dataset, 360B tokens |
| Is it open? | No β API only, internals secret | Yes β full source + model on GitHub & HuggingFace |
| Design goal | Speed + calibrated confidence for software decisions | Speed + tiny memory for on-device automation |
The overlap (this is why people get confused)
Both Jev and Needle:
- return structured values software can use directly (not just chat text),
- are obsessed with speed and efficiency over raw "smartness",
- both claim they can't hallucinate their structured outputs,
- both are aimed at automating software, not chatting with humans.
So at a product level they're chasing the same big idea: "stop using a giant chat model to answer tiny yes/no questions inside your code."
Where they genuinely diverge
- Jev gives up text generation entirely. It answers questions (is this urgent?) with a probability. It is, at its heart, a very smart classifier with calibrated confidence.
- Needle still generates. It produces tool calls as text tokens, it just makes that fast enough and small enough to run on a Raspberry Pi.
And the solutions are different too. Jev's architecture is a secret. Needle's is fully documented: Simple Attention Networks, no feed-forward layers, Hashed N-gram "engram" memory, Hadamard (Kronecker-factored) MLPs, monarch matrices, and 2-bit quantization. Those are Needle's design choices, not Jev's β and Jev's own choices are unpublished.
Honest takeaway: they're siblings in the "automation models" family, but different problems β different solutions. Jev = fast calibrated decision on structured state. Needle = tiny tool-calling + extraction model that runs anywhere.
Part 2 β Needle's architecture (this one IS fully documented)
Because Needle is open source, I can tell you exactly how it works. This is real, from cactuscompute.com/needle, the GitHub repo, and the Cactus engineering blog.
2.1 The headline: it's a transformer WITHOUT the expensive parts
Normal LLMs spend most of their FLOPs on two things: the feed-forward (MLP) layer and large vocabularies. Needle tackles both:
- No FFN (or almost none). A normal transformer's MLP is ~2/3 of its weights. Needle replaces it with a Hadamard/Kronecker-factored mixer β roughly 25.6K params per layer instead of 4.7M (per their blog). It burns ~1/5 the compute per token.
- Tiny vocabulary. 8,192 SentencePiece BPE tokens (normal models have 50k+). Smaller head, less memory, faster decoding.
2.2 The pieces (from the Figure 1 diagram on the site)
Reading the official diagram, the architecture has these components:
- Embedding β 8,192 Γ 768, tied (shared) with the unembedding. Input text β vectors.
- "mHC lane read" β a multi-head channel read: 4 residual lanes, combining
u = Ξ£β Ο(a Οβα΅ xΜ + b) Xβ(a gated weighted sum of the lane states). - Engram fusion β hashed 2- and 3-gram memory. This is a fast, learnable lookup:
x β x + Ο(β¨xΜ, kΜβ©/βd)Β·v, with 18,432 slots, read by gather. Think of it as a small in-model "memory table" that stores frequent patterns. - GQA attention + RoPE + QK-norm β standard grouped-query attention with rotary positions, window 1024, with global attention at layers 4, 9, 14, 19.
- Monarch Hadamard FFN β the FFN replacement:
y = DβMβDβMβ silu(Dβc(x)MβDβx + b), where eachMα΅’ = Aα΅’ β Bα΅’is a Kronecker product of 32Γ32 matrices. Cheap + surprisingly expressive. - mHC lane write β
Xβ² = P X + 2Ο(h) β y, withP = Sinkhorn(A)(a learned, doubly-stochastic projection). - ZCRMSNorm + tied unembed + byte-level grammar decoder that guarantees exact function calls (a constrained-decoding step β "can't hallucinate" the JSON).
2.3 "Intelligence Laddering" β one model, many sizes
This is the clever product trick: the 20-layer model is trained once, and any prefix of layers is its own smaller model (2 layers up to 20). Blocks 0 and 19 are always kept. So you pick "how smart" from the same weights:
- 2-layer subnetwork β 9 MB, for the tiniest devices
- 20-layer β 29 MB, full 121M params
And they claim fine-tuning just 4L on downstream tool-calling can match DeepSeek V4
Flash for that task. That's a big claim but it's testable against their FineTune setup.
2.4 The numbers they publish (needle site)
- Speed: 400β4,000 tok/s decode, 1β10k tok/s prefill, on a Raspberry Pi 5
- Size: 8β29 MB (CQ2 2-bit quantized), engine under 1 MB
- Accuracy: beats models 10Γ its size on mobile tool-calling; 2β3Γ on extraction
- Training: 360B tokens of proprietary structured data
These are their claims on their evals β I'm reporting them, not independently verifying. That's the honest framing.
Part 3 β what's actually proven about Jev
Now the project that's not open. Let me carefully separate fact from guess.
3.1 Proven (from the announcement, docs, and demos)
Source: TypeSafe intro blog, TypesSafe docs, LangChain's guide.
- It's a "System One Model" β intentionally named after Kahneman's fast, intuitive System 1 thinking (vs. slow "System 2" reasoning models).
- It does NOT generate text. It evaluates a
stateand answers typedquestions. - Input: a
state(string, JSON object, or array) + a set of questions. Each question has atype. - Three question types:
- noul β a yes/no statement, returns the probability it's true.
- choice β pick among options, returns a probability distribution.
- score β rate against ordered levels, returns a continuous score + distribution.
- Parallel: all questions on one state are answered in one shot, off the same state.
- Calibrated confidence: higher confidence βΆ more likely correct (they train for this).
- Speed: 70β500 ms end-to-end; they claim 40β200Γ faster / 400Γ cheaper than LLMs on decision-shaped queries.
- Can't hallucinate: outputs are constrained to the predefined schema, so they claim it "can't make type errors."
- Training method called RLCD (Reinforcement Learning for Calibrated Decisions).
3.2 NOT public (therefore we refuse to guess in the main narrative)
- The model architecture (transformer? MLP? what kind of heads?)
- Parameter count
- How the parallel sampler actually works internally
- What "RLCD" concretely does under the hood
This is important: several "Explain Jev" pieces on the internet invent an architecture. I'm deliberately not doing that. What the community does know, from reverse-engineering by building compatible clones, is in 3.3.
3.3 What the community learned by building Jev-compatible clones
This is real, reproducible knowledge from people who built servers that answer the Jev wire-protocol:
- OpenJev (Razorback16): a Jev-compatible server running on DiffusionGemma 26B. Key insight: to get probabilities like Jev, they use a "structured read" β a diffusion model's denoising pass in read-only mode. OpenJev needed unmerged custom vLLM extensions to do it.
- LocalJev (GitHub Next): a local
Jev-compatible API in TypeScript. Their key finding: you can translate
state+ typed questions into a classification prompt, ask an LLM for a JSON probability, validate, and normalize into the Jev response shape. But β and this is the honest caveat β those probabilities are the model's self-reported numbers, not direct logits, so calibration isn't guaranteed. - Bespoke Nimble (madiator): a
Jev-like model from a LoRA fine-tune of Qwen3.5-9B using synthetic contrastive data
- constrained decoding. Reported 66% β 90% on their eval vs. 93% for Jev itself.
What the clones tell us (verified): to reproduce Jev's behavior you need (a) a way to get probabilities efficiently (diffusion read / logits / prompted JSON), (b) typed output heads, (c) calibration, and (d) parallel evaluation. But none of this proves what Jev's actual weights look like. It's reconstruction, not revelation.
Part 4 β LET'S BUILD ONE FROM SCRATCH
Now the fun part. I built a working, naive "Jev-style" model. This is my design β for Jev's real architecture, see Part 3.2 (unknown). What I do deliberately copy from the proven interface:
- one
state, many typedquestions - answered in parallel (one encoder pass over the state)
- outputs are probabilities
- trained with a proper scoring rule (optimizing calibration)
- no text generation at all
Let me show you the whole thing, simplified, and then give you the real code.
4.1 The mental model
Think of it as: "read a page of testimony once, then answer every question on a worksheet from memory of that same reading."
state ("the article") questions ("is urgent?", "which topic?", "score anger 0-1")
β β
βΌ βΌ
βββββββββββββββββ βββββββββββββββββββ
β Encoder β (ONE pass) β Encoder β
β (transformer)β β (shared weights)β
βββββββββ¬ββββββββ ββββββββββ¬βββββββββ
β pooled vector h_state β h_question
βββββββββββββββββ¬βββββββββββββββββββ
βΌ
fuse: concat β MLP
β
βββββββββββββββββΌββββββββββββββββ
βΌ βΌ βΌ
noul head choice head score head
(sigmoid) (softmax over K) (logistic β range)
β β β
P(true) distribution score in [lo,hi]
β β β
βββββββββββββββββ΄ββββββββββββββββ
βΌ
all answers at once, from ONE state read
The key word is "ONE". We encode the state exactly once. Every question reuses that pooled state vector β that's the parallelism, and it's real, not a marketing trick.
Here is the actual diagram of our build (rendered from the code):
4.1a The input and the output (concretely)
So you can see exactly what goes in and what comes out, here is a real request we ran on our trained model (full transcript in the screenshot at the end of this section):
INPUT β one state, three typed questions:
{
"state": "I've been trying to connect my Stripe account for 3 days,
it keeps failing with a 403. I'm losing sales and my manager
is anxious. Please help ASAP.",
"questions": {
"is_urgent": { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity." },
"department": { "type": "choice", "instructions": "Which team should handle this?",
"options": ["billing", "technical", "sales"] },
"frustration":{"type": "score", "instructions": "How frustrated does the customer appear? (0=civil, 1=angry)" }
}
}
OUTPUT β one typed, probabilistic answer per question (no text generation):
{
"is_urgent": { "type": "noul", "noul": 0.6742, "is_true": true },
"department": { "type": "choice", "distribution": {"billing": 0.0, "technical": 1.0, "sales": 0.0},
"argmax": "technical", "confidence": 1.0 },
"frustration": { "type": "score", "score": 0.518 }
}
Every field is either a probability (noul), a probability distribution + confidence (choice), or a bounded number (score). No tokens are generated anywhere.
Screenshot of the real run (this is actual terminal output from our model):
4.2 Three heads for three question types
| Type | Head | Math | Output |
|---|---|---|---|
| noul | 1 logit β sigmoid | p = Ο(wΒ·h) |
probability (0β1) |
| choice | K logits β softmax | p_k = softmax(WΒ·h)_k |
distribution + confidence |
| score | 1 logit, logistic-mapped | s = lo + (hiβlo)Β·Ο(wΒ·h) |
bounded real value |
Because a decision model reads the whole context, we use no causal mask β every head can attend to the entire state. (A generative LLM can't do this; it must read left to right. This is a real, structural advantage of "decision" models.)
4.3 Why is it fast? (the honest reasons)
- No autoregressive generation. No
for token in .... One forward pass, done. Generation is where LLMs spend most of their latency. - No output tokens to pay for. Every "answer" is a handful of numbers.
- One state pass for N questions. Adding questions doesn't re-read the state.
- Small. Our toy is ~3M params; even real systems are far smaller than a frontier LLM.
4.4 Calibration β the piece that makes it "honest"
A model that guesses "urgent = 99%" but is only right 60% of the time is useless for
automation. So we train with a proper scoring rule (Brier score for noul β you're
penalized for confidence that doesn't match reality), and then we do temperature
scaling: search for a single scalar T that makes predicted probabilities match actual
accuracy on a held-out set. That's our (small) implementation of "calibrated decisions."
Part 5 β the actual code
All code is in this repo. Here are the three files that matter, then the full source.
jev_toy/model.pyβ the architecture (encoder + 3 heads)jev_toy/data.pyβ turns real datasets into (state, question, type, target) pairsjev_toy/train.pyβ training loop with parallel, per-head proper-scoring lossesjev_toy/eval.pyβ accuracy, Brier, expected-calibration-error (ECE), temperature scalejev_toy/serve.pyβ a mini "Jev API": one state, many questions, all answered in parallel
model.py (simplified)
class SystemOneModel(nn.Module):
def __init__(self, cfg):
self.encoder = _Encoder(cfg) # shared transformer encoder
self.noul_head = nn.Linear(cfg.d_model, 1)
self.choice_head = nn.Linear(cfg.d_model, cfg.num_choice_heads)
self.score_head = nn.Linear(cfg.d_model, 1)
def encode_state(self, s_ids, s_mask):
return self.encoder(s_ids, s_mask) # ONE pass over the state
def answer(self, h_state, q_ids, q_mask):
h_q = self.encoder(q_ids, q_mask) # encode each question
h = F.gelu(self.merge(torch.cat([h_state, h_q], -1)))
return {
"noul": self.noul_head(h),
"choice": self.choice_head(h),
"score": self.score_head(h),
}
The data (real, public datasets)
We map the three question types onto real Hugging Face datasets:
| Question type | Dataset | Task we phrase |
|---|---|---|
| choice | AG News (4 classes) | "which topic is this article?" |
| noul | BoolQ | "is this yes/no statement true?" |
| noul | SST-2 | "is this review positive?" |
(Toy scale: 1,500 examples each, subsampled so a laptop trains in minutes. The code scales to the full datasets and to GPU as-is.)
train.py (simplified β the parallel trick)
for batch in batches:
# ONE forward pass encodes all states in the batch
h_state = model.encode_state(state_ids, state_mask)
# every question (of any type) answered off that shared state
logits = model.answer(h_state, question_ids, question_mask)
loss = proper_scoring_loss(logits, types, targets) # Brier + CE + MSE
loss.backward(); opt.step()
Part 6 β real results (ran on my machine, no GPU)
This section is honest numbers from a real run β 8-core laptop CPU, 15 GB RAM, ~3.1M parameter model, 3 epochs, word-level vocab of ~20k.
| Question type | Dataset | Eval rows | Accuracy | Brier | ECE |
|---|---|---|---|---|---|
| noul | BoolQ + SST-2 | 2,000 | 59.7% | 0.236 | 0.0265 |
| choice | AG News (4-way) | 1,000 | 75.9% | 0.331 | β |
Training loss: 1.55 β 1.27 β 1.00 across three epochs. Post-hoc temperature scaling picked T = 0.8 (ECE β 0.0254).
Read these numbers honestly. Accuracy is low β that's expected and fine: ~3M params, a small slice of each dataset (1,500 of 120k AG News, 25% of BoolQ, 2% of SST-2), only 3 epochs, CPU-only. The interesting result is the calibration: an ECE of ~0.027 means the model's stated confidence genuinely matches how often it's right, within ~3 percentage points. That's the whole architectural point of a "System One" model β not raw smartness, but honest smartness a program can safely branch on.
And the serving demo returns the exact Jev response shape from one state + three parallel
questions (full transcript in RESULTS.md).
Does our architecture actually predict in parallel β or is it just a fine-tuned transformer faking it? This is the honest question, so let's answer it with a measurement, not an adjective. We ran the full answer path on one state while increasing the number of questions, on CPU:
| Questions asked | Full answer time (ms) |
|---|---|
| 1 | 27.3 |
| 3 | 35.7 |
| 6 | 46.3 |
| 12 | 101.0 |
If we were re-reading the state for every question (the naive fine-tuned-transformer approach), going from 1 β 12 questions would cost ~12Γ the time. Instead it costs ~3.7Γ. The state encoder runs exactly once; the growth comes only from the cheap question-batch pass. That is the parallel property Jev claims, made measurable in our build. (This is a CPU number β the structure is what matters, not the milliseconds.)
How it differs from "a transformer fine-tuned to act like Jev": a normal fine-tuned LLM generates an answer token-by-token (slow, sequential, can drift off-schema). Ours has no generation loop at all β the encoder runs, heads fire, done. That's the structural difference, not just a speed trick.
The value of this build is architectural β you can hold it all in your head, watch it train, and share one encoder across three question types. Full data + deeper encoder + GPU is the same code with bigger numbers.
Part 7 β the honest limitations (no hype)
- This is not Jev. It's a toy reconstruction of Jev's proven interface, with my own internals. Nothing here claims to be TypeSafe's model.
- No frontier accuracy. Toy model + tiny data + CPU = learning architecture, not SOTA. The community clones (OpenJev on DiffusionGemma 26B) are the real "reproduce the capability" efforts.
- Calibration at toy scale is limited. Post-hoc temperature scaling helps, but a model with 60% accuracy can only be calibrated, not correct. Calibration is about honesty, not magic.
- The
choiceconfidence formula (max prob minus uniform margin) is a community convention from the Turing Post parse of the TypeSafe adapter β it's what they use in their open adapter, but treat it as a heuristic. - Needle and Jev are different. Don't repeat the internet's mistake of treating them as the same thing.
Part 8 β what to read next (all real sources)
Jev / TypeSafe
- Announcement: https://typesafe.ai/blog/introducing-system-one-models-and-jev
- Docs (state, primitives): https://docs.typesafe.ai/concepts/state
- LangChain's guide: https://www.langchain.com/blog/building-a-harness-with-jev
- Turing Post analysis: https://www.turingpost.com/p/what-is-jev-rlcd
- Clones: OpenJev, LocalJev, awesome-typesafe
Needle / Cactus
- Model & sandbox: https://cactuscompute.com/needle
- Source: https://github.com/cactus-compute/needle
- HuggingFace: https://huggingface.co/Cactus-Compute/needle
- Blog (Hadamard MLP, .cact format): https://cactuscompute.com/blog
Part 9 β quick answers to your original question
Does Jev (TypeScript) and Needle solve the same problem?
No. Jev = fast, calibrated, typed decisions from a proprietary cloud model. Needle = tiny tool-calling + structured extraction on-device model. They share the goal of "structured output for software automation" but attack different problems with different (Needle's is public; Jev's is secret) solutions.
Can we build Jev from scratch?
We can build a faithful reproduction of Jev's proven interface from scratch β which is exactly what this article + code does. We cannot build Jev's actual internals from scratch because they were never published. That distinction matters, and I've tried to hold it throughout.
Built by Hermes for the "build the Jev from scratch" deep-dive. Grounded in linked sources and a real CPU training run. Toy scale, honest framing, no invented research.

