azharmo commited on
Commit
f48dabc
·
verified ·
1 Parent(s): 877a9be

Upload ARTICLE.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. ARTICLE.md +64 -0
ARTICLE.md CHANGED
@@ -247,6 +247,47 @@ worksheet from memory of that same reading."*
247
  **The key word is "ONE".** We encode the state exactly once. Every question reuses that
248
  pooled state vector — that's the parallelism, and it's real, not a marketing trick.
249
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
250
  ### 4.2 Three heads for three question types
251
 
252
  | Type | Head | Math | Output |
@@ -360,6 +401,29 @@ but *honest* smartness a program can safely branch on.
360
  And the serving demo returns the exact Jev response shape from one state + three parallel
361
  questions (full transcript in `RESULTS.md`).
362
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
363
  The value of this build is architectural — you can hold it all in your head, watch it
364
  train, and share one encoder across three question types. Full data + deeper encoder + GPU
365
  is the same code with bigger numbers.
 
247
  **The key word is "ONE".** We encode the state exactly once. Every question reuses that
248
  pooled state vector — that's the parallelism, and it's real, not a marketing trick.
249
 
250
+ Here is the actual diagram of our build (rendered from the code):
251
+
252
+ ![architecture](ARCHITECTURE.png)
253
+
254
+ ### 4.1a The input and the output (concretely)
255
+
256
+ So you can see exactly what goes in and what comes out, here is a real request we ran on
257
+ our trained model (full transcript in the screenshot at the end of this section):
258
+
259
+ **INPUT — one `state`, three typed `questions`:**
260
+ ```json
261
+ {
262
+ "state": "I've been trying to connect my Stripe account for 3 days,
263
+ it keeps failing with a 403. I'm losing sales and my manager
264
+ is anxious. Please help ASAP.",
265
+ "questions": {
266
+ "is_urgent": { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity." },
267
+ "department": { "type": "choice", "instructions": "Which team should handle this?",
268
+ "options": ["billing", "technical", "sales"] },
269
+ "frustration":{"type": "score", "instructions": "How frustrated does the customer appear? (0=civil, 1=angry)" }
270
+ }
271
+ }
272
+ ```
273
+
274
+ **OUTPUT — one typed, probabilistic answer per question (no text generation):**
275
+ ```json
276
+ {
277
+ "is_urgent": { "type": "noul", "noul": 0.6742, "is_true": true },
278
+ "department": { "type": "choice", "distribution": {"billing": 0.0, "technical": 1.0, "sales": 0.0},
279
+ "argmax": "technical", "confidence": 1.0 },
280
+ "frustration": { "type": "score", "score": 0.518 }
281
+ }
282
+ ```
283
+
284
+ Every field is either a probability (noul), a probability distribution + confidence
285
+ (choice), or a bounded number (score). No tokens are generated anywhere.
286
+
287
+ **Screenshot of the real run** (this is actual terminal output from our model):
288
+
289
+ ![serve screenshot](SCREENSHOT_serve.png)
290
+
291
  ### 4.2 Three heads for three question types
292
 
293
  | Type | Head | Math | Output |
 
401
  And the serving demo returns the exact Jev response shape from one state + three parallel
402
  questions (full transcript in `RESULTS.md`).
403
 
404
+ **Does our architecture *actually* predict in parallel — or is it just a fine-tuned
405
+ transformer faking it?** This is the honest question, so let's answer it with a
406
+ measurement, not an adjective. We ran the full answer path on one state while increasing
407
+ the number of questions, on CPU:
408
+
409
+ | Questions asked | Full answer time (ms) |
410
+ |---|---|
411
+ | 1 | 27.3 |
412
+ | 3 | 35.7 |
413
+ | 6 | 46.3 |
414
+ | 12 | 101.0 |
415
+
416
+ If we were *re-reading the state for every question* (the naive fine-tuned-transformer
417
+ approach), going from 1 → 12 questions would cost **~12×** the time. Instead it costs
418
+ **~3.7×**. The state encoder runs exactly once; the growth comes only from the cheap
419
+ question-batch pass. That is the parallel property Jev claims, made measurable in our
420
+ build. (This is a CPU number — the *structure* is what matters, not the milliseconds.)
421
+
422
+ How it differs from "a transformer fine-tuned to act like Jev": a normal fine-tuned LLM
423
+ *generates* an answer token-by-token (slow, sequential, can drift off-schema). Ours has
424
+ **no generation loop at all** — the encoder runs, heads fire, done. That's the structural
425
+ difference, not just a speed trick.
426
+
427
  The value of this build is architectural — you can hold it all in your head, watch it
428
  train, and share one encoder across three question types. Full data + deeper encoder + GPU
429
  is the same code with bigger numbers.