azharmo's picture
Upload ARTICLE.md with huggingface_hub
f48dabc verified
|
Raw History Blame
24.3 kB

Let's Build a "Jev" From Scratch (Simple English Edition)

An honest, from-the-ground-up walkthrough of a System One decision model β€” and how it compares to Cactus's Needle.

Read this first β€” the truth up front. The real Jev is made by TypeSafe AI (announced Sept 15, 2026). Its actual model internals are not public β€” it's a proprietary, invite-only API. Nobody outside TypeSafe knows exactly how Jev is built.

So this article does not pretend to show you "the real Jev". That would be a lie. Instead I do three honest things:

  1. Tell you exactly what about Jev is publicly proven (the interface, the claims).
  2. Compare it to Needle (Cactus Compute), which is fully open source.
  3. Build my own toy "Jev-style" model from scratch β€” our design, clearly labeled, trained on real public datasets on a laptop with no GPU, and measured honestly.

Everything I claim here is either linked to a source or was produced by code that actually ran on my machine. No fake numbers, no invented research.


Part 0 β€” who is talking to you

Hi. I read the TypeSafe announcement, the Cactus docs, the community clones (OpenJev, LocalJev, Bespoke Nimble), and then actually built and trained a model. This is me explaining what I found and what I built, the way I'd explain it to a friend over coffee. If a sentence sounds like a Wikipedia fact, it has a source link. If it sounds like my opinion, it's mine.


Part 1 β€” the big question: do Jev and Needle solve the same problem?

This is what you asked, so let's answer it directly.

Short answer: no, they're different projects that overlap in one useful place.

Let me be precise, because people are conflating them everywhere right now.

What each one is

Jev (TypeSafe AI) Needle (Cactus Compute)
What it does Makes fast, typed decisions (is this urgent? pick a category, score from 1–5) Calls tools / extracts structure β€” a tiny on-device "automation" model
Output Probabilities + typed values (noul / choice / score). No text generation. Structured JSON (tool calls, extractions) via token-by-token generation
Where it runs Cloud API (early access), 70–500 ms On-device β€” phones, Raspberry Pi, browsers (WASM), a 9–29 MB file
Size Unknown (proprietary) 9–121M params, 2-bit quantized, 8–29 MB
Training RLCD (Reinforcement Learning for Calibrated Decisions) β€” proprietary Proprietary structured dataset, 360B tokens
Is it open? No β€” API only, internals secret Yes β€” full source + model on GitHub & HuggingFace
Design goal Speed + calibrated confidence for software decisions Speed + tiny memory for on-device automation

The overlap (this is why people get confused)

Both Jev and Needle:

  • return structured values software can use directly (not just chat text),
  • are obsessed with speed and efficiency over raw "smartness",
  • both claim they can't hallucinate their structured outputs,
  • both are aimed at automating software, not chatting with humans.

So at a product level they're chasing the same big idea: "stop using a giant chat model to answer tiny yes/no questions inside your code."

Where they genuinely diverge

  • Jev gives up text generation entirely. It answers questions (is this urgent?) with a probability. It is, at its heart, a very smart classifier with calibrated confidence.
  • Needle still generates. It produces tool calls as text tokens, it just makes that fast enough and small enough to run on a Raspberry Pi.

And the solutions are different too. Jev's architecture is a secret. Needle's is fully documented: Simple Attention Networks, no feed-forward layers, Hashed N-gram "engram" memory, Hadamard (Kronecker-factored) MLPs, monarch matrices, and 2-bit quantization. Those are Needle's design choices, not Jev's β€” and Jev's own choices are unpublished.

Honest takeaway: they're siblings in the "automation models" family, but different problems β†’ different solutions. Jev = fast calibrated decision on structured state. Needle = tiny tool-calling + extraction model that runs anywhere.


Part 2 β€” Needle's architecture (this one IS fully documented)

Because Needle is open source, I can tell you exactly how it works. This is real, from cactuscompute.com/needle, the GitHub repo, and the Cactus engineering blog.

2.1 The headline: it's a transformer WITHOUT the expensive parts

Normal LLMs spend most of their FLOPs on two things: the feed-forward (MLP) layer and large vocabularies. Needle tackles both:

  • No FFN (or almost none). A normal transformer's MLP is ~2/3 of its weights. Needle replaces it with a Hadamard/Kronecker-factored mixer β€” roughly 25.6K params per layer instead of 4.7M (per their blog). It burns ~1/5 the compute per token.
  • Tiny vocabulary. 8,192 SentencePiece BPE tokens (normal models have 50k+). Smaller head, less memory, faster decoding.

2.2 The pieces (from the Figure 1 diagram on the site)

Reading the official diagram, the architecture has these components:

  1. Embedding β€” 8,192 Γ— 768, tied (shared) with the unembedding. Input text β†’ vectors.
  2. "mHC lane read" β€” a multi-head channel read: 4 residual lanes, combining u = Ξ£β‚™ Οƒ(a Ο†β‚™α΅€ xΜ‚ + b) Xβ‚™ (a gated weighted sum of the lane states).
  3. Engram fusion β€” hashed 2- and 3-gram memory. This is a fast, learnable lookup: x ← x + Οƒ(⟨xΜ‚, kΜ‚βŸ©/√d)Β·v, with 18,432 slots, read by gather. Think of it as a small in-model "memory table" that stores frequent patterns.
  4. GQA attention + RoPE + QK-norm β€” standard grouped-query attention with rotary positions, window 1024, with global attention at layers 4, 9, 14, 19.
  5. Monarch Hadamard FFN β€” the FFN replacement: y = Dβ‚„M₃D₃Mβ‚‚ silu(Dβ‚‚c(x)M₁D₁x + b), where each Mα΅’ = Aα΅’ βŠ— Bα΅’ is a Kronecker product of 32Γ—32 matrices. Cheap + surprisingly expressive.
  6. mHC lane write β€” Xβ€² = P X + 2Οƒ(h) βŠ— y, with P = Sinkhorn(A) (a learned, doubly-stochastic projection).
  7. ZCRMSNorm + tied unembed + byte-level grammar decoder that guarantees exact function calls (a constrained-decoding step β†’ "can't hallucinate" the JSON).

2.3 "Intelligence Laddering" β€” one model, many sizes

This is the clever product trick: the 20-layer model is trained once, and any prefix of layers is its own smaller model (2 layers up to 20). Blocks 0 and 19 are always kept. So you pick "how smart" from the same weights:

  • 2-layer subnetwork β‰ˆ 9 MB, for the tiniest devices
  • 20-layer β‰ˆ 29 MB, full 121M params

And they claim fine-tuning just 4L on downstream tool-calling can match DeepSeek V4 Flash for that task. That's a big claim but it's testable against their FineTune setup.

2.4 The numbers they publish (needle site)

  • Speed: 400–4,000 tok/s decode, 1–10k tok/s prefill, on a Raspberry Pi 5
  • Size: 8–29 MB (CQ2 2-bit quantized), engine under 1 MB
  • Accuracy: beats models 10Γ— its size on mobile tool-calling; 2–3Γ— on extraction
  • Training: 360B tokens of proprietary structured data

These are their claims on their evals β€” I'm reporting them, not independently verifying. That's the honest framing.


Part 3 β€” what's actually proven about Jev

Now the project that's not open. Let me carefully separate fact from guess.

3.1 Proven (from the announcement, docs, and demos)

Source: TypeSafe intro blog, TypesSafe docs, LangChain's guide.

  • It's a "System One Model" β€” intentionally named after Kahneman's fast, intuitive System 1 thinking (vs. slow "System 2" reasoning models).
  • It does NOT generate text. It evaluates a state and answers typed questions.
  • Input: a state (string, JSON object, or array) + a set of questions. Each question has a type.
  • Three question types:
    • noul β€” a yes/no statement, returns the probability it's true.
    • choice β€” pick among options, returns a probability distribution.
    • score β€” rate against ordered levels, returns a continuous score + distribution.
  • Parallel: all questions on one state are answered in one shot, off the same state.
  • Calibrated confidence: higher confidence ⟢ more likely correct (they train for this).
  • Speed: 70–500 ms end-to-end; they claim 40–200Γ— faster / 400Γ— cheaper than LLMs on decision-shaped queries.
  • Can't hallucinate: outputs are constrained to the predefined schema, so they claim it "can't make type errors."
  • Training method called RLCD (Reinforcement Learning for Calibrated Decisions).

3.2 NOT public (therefore we refuse to guess in the main narrative)

  • The model architecture (transformer? MLP? what kind of heads?)
  • Parameter count
  • How the parallel sampler actually works internally
  • What "RLCD" concretely does under the hood

This is important: several "Explain Jev" pieces on the internet invent an architecture. I'm deliberately not doing that. What the community does know, from reverse-engineering by building compatible clones, is in 3.3.

3.3 What the community learned by building Jev-compatible clones

This is real, reproducible knowledge from people who built servers that answer the Jev wire-protocol:

  • OpenJev (Razorback16): a Jev-compatible server running on DiffusionGemma 26B. Key insight: to get probabilities like Jev, they use a "structured read" β€” a diffusion model's denoising pass in read-only mode. OpenJev needed unmerged custom vLLM extensions to do it.
  • LocalJev (GitHub Next): a local Jev-compatible API in TypeScript. Their key finding: you can translate state + typed questions into a classification prompt, ask an LLM for a JSON probability, validate, and normalize into the Jev response shape. But β€” and this is the honest caveat β€” those probabilities are the model's self-reported numbers, not direct logits, so calibration isn't guaranteed.
  • Bespoke Nimble (madiator): a Jev-like model from a LoRA fine-tune of Qwen3.5-9B using synthetic contrastive data
    • constrained decoding. Reported 66% β†’ 90% on their eval vs. 93% for Jev itself.

What the clones tell us (verified): to reproduce Jev's behavior you need (a) a way to get probabilities efficiently (diffusion read / logits / prompted JSON), (b) typed output heads, (c) calibration, and (d) parallel evaluation. But none of this proves what Jev's actual weights look like. It's reconstruction, not revelation.


Part 4 β€” LET'S BUILD ONE FROM SCRATCH

Now the fun part. I built a working, naive "Jev-style" model. This is my design β€” for Jev's real architecture, see Part 3.2 (unknown). What I do deliberately copy from the proven interface:

  • one state, many typed questions
  • answered in parallel (one encoder pass over the state)
  • outputs are probabilities
  • trained with a proper scoring rule (optimizing calibration)
  • no text generation at all

Let me show you the whole thing, simplified, and then give you the real code.

4.1 The mental model

Think of it as: "read a page of testimony once, then answer every question on a worksheet from memory of that same reading."

    state ("the article")          questions  ("is urgent?", "which topic?", "score anger 0-1")
          β”‚                                   β”‚
          β–Ό                                   β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚    Encoder    β”‚  (ONE pass)     β”‚   Encoder       β”‚
  β”‚  (transformer)β”‚                 β”‚ (shared weights)β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜                 β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚ pooled vector h_state            β”‚ h_question
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β–Ό
                 fuse:  concat β†’ MLP
                          β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β–Ό               β–Ό               β–Ό
      noul head       choice head      score head
     (sigmoid)       (softmax over K)  (logistic β†’ range)
          β”‚               β”‚               β”‚
      P(true)        distribution    score in [lo,hi]
          β”‚               β”‚               β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                          β–Ό
              all answers at once, from ONE state read

The key word is "ONE". We encode the state exactly once. Every question reuses that pooled state vector β€” that's the parallelism, and it's real, not a marketing trick.

Here is the actual diagram of our build (rendered from the code):

architecture

4.1a The input and the output (concretely)

So you can see exactly what goes in and what comes out, here is a real request we ran on our trained model (full transcript in the screenshot at the end of this section):

INPUT β€” one state, three typed questions:

{
  "state": "I've been trying to connect my Stripe account for 3 days,
            it keeps failing with a 403. I'm losing sales and my manager
            is anxious. Please help ASAP.",
  "questions": {
    "is_urgent":  { "type": "noul",   "instructions": "The message conveys urgency or time-sensitivity." },
    "department": { "type": "choice", "instructions": "Which team should handle this?",
                    "options": ["billing", "technical", "sales"] },
    "frustration":{"type": "score",   "instructions": "How frustrated does the customer appear? (0=civil, 1=angry)" }
  }
}

OUTPUT β€” one typed, probabilistic answer per question (no text generation):

{
  "is_urgent":   { "type": "noul",   "noul": 0.6742, "is_true": true },
  "department":  { "type": "choice", "distribution": {"billing": 0.0, "technical": 1.0, "sales": 0.0},
                   "argmax": "technical", "confidence": 1.0 },
  "frustration": { "type": "score",  "score": 0.518 }
}

Every field is either a probability (noul), a probability distribution + confidence (choice), or a bounded number (score). No tokens are generated anywhere.

Screenshot of the real run (this is actual terminal output from our model):

serve screenshot

4.2 Three heads for three question types

Type Head Math Output
noul 1 logit β†’ sigmoid p = Οƒ(wΒ·h) probability (0–1)
choice K logits β†’ softmax p_k = softmax(WΒ·h)_k distribution + confidence
score 1 logit, logistic-mapped s = lo + (hiβˆ’lo)Β·Οƒ(wΒ·h) bounded real value

Because a decision model reads the whole context, we use no causal mask β€” every head can attend to the entire state. (A generative LLM can't do this; it must read left to right. This is a real, structural advantage of "decision" models.)

4.3 Why is it fast? (the honest reasons)

  1. No autoregressive generation. No for token in .... One forward pass, done. Generation is where LLMs spend most of their latency.
  2. No output tokens to pay for. Every "answer" is a handful of numbers.
  3. One state pass for N questions. Adding questions doesn't re-read the state.
  4. Small. Our toy is ~3M params; even real systems are far smaller than a frontier LLM.

4.4 Calibration β€” the piece that makes it "honest"

A model that guesses "urgent = 99%" but is only right 60% of the time is useless for automation. So we train with a proper scoring rule (Brier score for noul β€” you're penalized for confidence that doesn't match reality), and then we do temperature scaling: search for a single scalar T that makes predicted probabilities match actual accuracy on a held-out set. That's our (small) implementation of "calibrated decisions."


Part 5 β€” the actual code

All code is in this repo. Here are the three files that matter, then the full source.

  • jev_toy/model.py β€” the architecture (encoder + 3 heads)
  • jev_toy/data.py β€” turns real datasets into (state, question, type, target) pairs
  • jev_toy/train.py β€” training loop with parallel, per-head proper-scoring losses
  • jev_toy/eval.py β€” accuracy, Brier, expected-calibration-error (ECE), temperature scale
  • jev_toy/serve.py β€” a mini "Jev API": one state, many questions, all answered in parallel

model.py (simplified)

class SystemOneModel(nn.Module):
    def __init__(self, cfg):
        self.encoder = _Encoder(cfg)              # shared transformer encoder
        self.noul_head   = nn.Linear(cfg.d_model, 1)
        self.choice_head = nn.Linear(cfg.d_model, cfg.num_choice_heads)
        self.score_head  = nn.Linear(cfg.d_model, 1)

    def encode_state(self, s_ids, s_mask):
        return self.encoder(s_ids, s_mask)         # ONE pass over the state

    def answer(self, h_state, q_ids, q_mask):
        h_q = self.encoder(q_ids, q_mask)          # encode each question
        h = F.gelu(self.merge(torch.cat([h_state, h_q], -1)))
        return {
            "noul":   self.noul_head(h),
            "choice": self.choice_head(h),
            "score":  self.score_head(h),
        }

The data (real, public datasets)

We map the three question types onto real Hugging Face datasets:

Question type Dataset Task we phrase
choice AG News (4 classes) "which topic is this article?"
noul BoolQ "is this yes/no statement true?"
noul SST-2 "is this review positive?"

(Toy scale: 1,500 examples each, subsampled so a laptop trains in minutes. The code scales to the full datasets and to GPU as-is.)

train.py (simplified β€” the parallel trick)

for batch in batches:
    # ONE forward pass encodes all states in the batch
    h_state = model.encode_state(state_ids, state_mask)
    # every question (of any type) answered off that shared state
    logits = model.answer(h_state, question_ids, question_mask)
    loss = proper_scoring_loss(logits, types, targets)   # Brier + CE + MSE
    loss.backward(); opt.step()

Part 6 β€” real results (ran on my machine, no GPU)

This section is honest numbers from a real run β€” 8-core laptop CPU, 15 GB RAM, ~3.1M parameter model, 3 epochs, word-level vocab of ~20k.

Question type Dataset Eval rows Accuracy Brier ECE
noul BoolQ + SST-2 2,000 59.7% 0.236 0.0265
choice AG News (4-way) 1,000 75.9% 0.331 β€”

Training loss: 1.55 β†’ 1.27 β†’ 1.00 across three epochs. Post-hoc temperature scaling picked T = 0.8 (ECE β†’ 0.0254).

Read these numbers honestly. Accuracy is low β€” that's expected and fine: ~3M params, a small slice of each dataset (1,500 of 120k AG News, 25% of BoolQ, 2% of SST-2), only 3 epochs, CPU-only. The interesting result is the calibration: an ECE of ~0.027 means the model's stated confidence genuinely matches how often it's right, within ~3 percentage points. That's the whole architectural point of a "System One" model β€” not raw smartness, but honest smartness a program can safely branch on.

And the serving demo returns the exact Jev response shape from one state + three parallel questions (full transcript in RESULTS.md).

Does our architecture actually predict in parallel β€” or is it just a fine-tuned transformer faking it? This is the honest question, so let's answer it with a measurement, not an adjective. We ran the full answer path on one state while increasing the number of questions, on CPU:

Questions asked Full answer time (ms)
1 27.3
3 35.7
6 46.3
12 101.0

If we were re-reading the state for every question (the naive fine-tuned-transformer approach), going from 1 β†’ 12 questions would cost ~12Γ— the time. Instead it costs ~3.7Γ—. The state encoder runs exactly once; the growth comes only from the cheap question-batch pass. That is the parallel property Jev claims, made measurable in our build. (This is a CPU number β€” the structure is what matters, not the milliseconds.)

How it differs from "a transformer fine-tuned to act like Jev": a normal fine-tuned LLM generates an answer token-by-token (slow, sequential, can drift off-schema). Ours has no generation loop at all β€” the encoder runs, heads fire, done. That's the structural difference, not just a speed trick.

The value of this build is architectural β€” you can hold it all in your head, watch it train, and share one encoder across three question types. Full data + deeper encoder + GPU is the same code with bigger numbers.


Part 7 β€” the honest limitations (no hype)

  1. This is not Jev. It's a toy reconstruction of Jev's proven interface, with my own internals. Nothing here claims to be TypeSafe's model.
  2. No frontier accuracy. Toy model + tiny data + CPU = learning architecture, not SOTA. The community clones (OpenJev on DiffusionGemma 26B) are the real "reproduce the capability" efforts.
  3. Calibration at toy scale is limited. Post-hoc temperature scaling helps, but a model with 60% accuracy can only be calibrated, not correct. Calibration is about honesty, not magic.
  4. The choice confidence formula (max prob minus uniform margin) is a community convention from the Turing Post parse of the TypeSafe adapter β€” it's what they use in their open adapter, but treat it as a heuristic.
  5. Needle and Jev are different. Don't repeat the internet's mistake of treating them as the same thing.

Part 8 β€” what to read next (all real sources)

Jev / TypeSafe

Needle / Cactus


Part 9 β€” quick answers to your original question

Does Jev (TypeScript) and Needle solve the same problem?

No. Jev = fast, calibrated, typed decisions from a proprietary cloud model. Needle = tiny tool-calling + structured extraction on-device model. They share the goal of "structured output for software automation" but attack different problems with different (Needle's is public; Jev's is secret) solutions.

Can we build Jev from scratch?

We can build a faithful reproduction of Jev's proven interface from scratch β€” which is exactly what this article + code does. We cannot build Jev's actual internals from scratch because they were never published. That distinction matters, and I've tried to hold it throughout.


Built by Hermes for the "build the Jev from scratch" deep-dive. Grounded in linked sources and a real CPU training run. Toy scale, honest framing, no invented research.