Upload ARTICLE.md with huggingface_hub
Browse files- ARTICLE.md +64 -0
ARTICLE.md
CHANGED
|
@@ -247,6 +247,47 @@ worksheet from memory of that same reading."*
|
|
| 247 |
**The key word is "ONE".** We encode the state exactly once. Every question reuses that
|
| 248 |
pooled state vector — that's the parallelism, and it's real, not a marketing trick.
|
| 249 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 250 |
### 4.2 Three heads for three question types
|
| 251 |
|
| 252 |
| Type | Head | Math | Output |
|
|
@@ -360,6 +401,29 @@ but *honest* smartness a program can safely branch on.
|
|
| 360 |
And the serving demo returns the exact Jev response shape from one state + three parallel
|
| 361 |
questions (full transcript in `RESULTS.md`).
|
| 362 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 363 |
The value of this build is architectural — you can hold it all in your head, watch it
|
| 364 |
train, and share one encoder across three question types. Full data + deeper encoder + GPU
|
| 365 |
is the same code with bigger numbers.
|
|
|
|
| 247 |
**The key word is "ONE".** We encode the state exactly once. Every question reuses that
|
| 248 |
pooled state vector — that's the parallelism, and it's real, not a marketing trick.
|
| 249 |
|
| 250 |
+
Here is the actual diagram of our build (rendered from the code):
|
| 251 |
+
|
| 252 |
+

|
| 253 |
+
|
| 254 |
+
### 4.1a The input and the output (concretely)
|
| 255 |
+
|
| 256 |
+
So you can see exactly what goes in and what comes out, here is a real request we ran on
|
| 257 |
+
our trained model (full transcript in the screenshot at the end of this section):
|
| 258 |
+
|
| 259 |
+
**INPUT — one `state`, three typed `questions`:**
|
| 260 |
+
```json
|
| 261 |
+
{
|
| 262 |
+
"state": "I've been trying to connect my Stripe account for 3 days,
|
| 263 |
+
it keeps failing with a 403. I'm losing sales and my manager
|
| 264 |
+
is anxious. Please help ASAP.",
|
| 265 |
+
"questions": {
|
| 266 |
+
"is_urgent": { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity." },
|
| 267 |
+
"department": { "type": "choice", "instructions": "Which team should handle this?",
|
| 268 |
+
"options": ["billing", "technical", "sales"] },
|
| 269 |
+
"frustration":{"type": "score", "instructions": "How frustrated does the customer appear? (0=civil, 1=angry)" }
|
| 270 |
+
}
|
| 271 |
+
}
|
| 272 |
+
```
|
| 273 |
+
|
| 274 |
+
**OUTPUT — one typed, probabilistic answer per question (no text generation):**
|
| 275 |
+
```json
|
| 276 |
+
{
|
| 277 |
+
"is_urgent": { "type": "noul", "noul": 0.6742, "is_true": true },
|
| 278 |
+
"department": { "type": "choice", "distribution": {"billing": 0.0, "technical": 1.0, "sales": 0.0},
|
| 279 |
+
"argmax": "technical", "confidence": 1.0 },
|
| 280 |
+
"frustration": { "type": "score", "score": 0.518 }
|
| 281 |
+
}
|
| 282 |
+
```
|
| 283 |
+
|
| 284 |
+
Every field is either a probability (noul), a probability distribution + confidence
|
| 285 |
+
(choice), or a bounded number (score). No tokens are generated anywhere.
|
| 286 |
+
|
| 287 |
+
**Screenshot of the real run** (this is actual terminal output from our model):
|
| 288 |
+
|
| 289 |
+

|
| 290 |
+
|
| 291 |
### 4.2 Three heads for three question types
|
| 292 |
|
| 293 |
| Type | Head | Math | Output |
|
|
|
|
| 401 |
And the serving demo returns the exact Jev response shape from one state + three parallel
|
| 402 |
questions (full transcript in `RESULTS.md`).
|
| 403 |
|
| 404 |
+
**Does our architecture *actually* predict in parallel — or is it just a fine-tuned
|
| 405 |
+
transformer faking it?** This is the honest question, so let's answer it with a
|
| 406 |
+
measurement, not an adjective. We ran the full answer path on one state while increasing
|
| 407 |
+
the number of questions, on CPU:
|
| 408 |
+
|
| 409 |
+
| Questions asked | Full answer time (ms) |
|
| 410 |
+
|---|---|
|
| 411 |
+
| 1 | 27.3 |
|
| 412 |
+
| 3 | 35.7 |
|
| 413 |
+
| 6 | 46.3 |
|
| 414 |
+
| 12 | 101.0 |
|
| 415 |
+
|
| 416 |
+
If we were *re-reading the state for every question* (the naive fine-tuned-transformer
|
| 417 |
+
approach), going from 1 → 12 questions would cost **~12×** the time. Instead it costs
|
| 418 |
+
**~3.7×**. The state encoder runs exactly once; the growth comes only from the cheap
|
| 419 |
+
question-batch pass. That is the parallel property Jev claims, made measurable in our
|
| 420 |
+
build. (This is a CPU number — the *structure* is what matters, not the milliseconds.)
|
| 421 |
+
|
| 422 |
+
How it differs from "a transformer fine-tuned to act like Jev": a normal fine-tuned LLM
|
| 423 |
+
*generates* an answer token-by-token (slow, sequential, can drift off-schema). Ours has
|
| 424 |
+
**no generation loop at all** — the encoder runs, heads fire, done. That's the structural
|
| 425 |
+
difference, not just a speed trick.
|
| 426 |
+
|
| 427 |
The value of this build is architectural — you can hold it all in your head, watch it
|
| 428 |
train, and share one encoder across three question types. Full data + deeper encoder + GPU
|
| 429 |
is the same code with bigger numbers.
|