|
Download source/RESEARCH_BRIEF.md from andyshu/opensysone: direct link, hf CLI and curl.
- Browser
- Download file 24.9 kB
-
https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/RESEARCH_BRIEF.md
- Command line
-
hf download hf://andyshu/opensysone@58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/RESEARCH_BRIEF.md
-
curl -L -o RESEARCH_BRIEF.md https://huggingface.co/andyshu/opensysone/resolve/58f289696f58962a8ec98293d7b1abf9fd0c6b8b/source/RESEARCH_BRIEF.md
24.9 kB
| # Handover: Build a Jev-like Decision Model on 3 DGX Sparks | |
| ## Objective | |
| Build and evaluate a small-scale reproduction of the core ideas behind TypeSafe.ai's Jev / “System One” model using approximately three NVIDIA DGX Sparks. | |
| The objective is **not** to reproduce TypeSafe's proprietary implementation exactly. The objective is to determine whether the following design can be made practical: | |
| > A pretrained Transformer is converted from an autoregressive text generator into a low-latency, zero-shot decision model that evaluates arbitrary natural-language questions and candidate outcomes in parallel and returns calibrated probability distributions directly. | |
| The primary success criterion is evidence that this architecture can outperform or materially improve upon ordinary prompt-and-generate classification in at least some combination of: | |
| - latency; | |
| - throughput; | |
| - probability calibration; | |
| - consistency; | |
| - cost per decision; | |
| - scaling to many independent questions over one shared context. | |
| Treat architecture claims about Jev as hypotheses unless independently verified. | |
| --- | |
| # 1. Available hardware | |
| Assume access to: | |
| - 3 × NVIDIA DGX Spark; | |
| - 128 GB unified memory per Spark; | |
| - GB10 GPU; | |
| - approximately 273 GB/s local memory bandwidth per node; | |
| - 200 Gbit/s ConnectX-7 networking between Sparks; | |
| - direct Spark-to-Spark topology if practical. | |
| Do not assume that the combined 384 GB behaves as unified GPU memory. | |
| Prefer models that fit entirely on one Spark for inference. | |
| Use multiple Sparks primarily for: | |
| - distributed fine-tuning; | |
| - synthetic data generation; | |
| - evaluation; | |
| - teacher inference; | |
| - parallel experiments. | |
| Avoid tensor-parallel inference across Sparks unless necessary because network bandwidth is much lower than local memory bandwidth. | |
| --- | |
| # 2. Core hypothesis | |
| The highest-priority architecture to test is: | |
| ```text | |
| shared state/context | |
| | | |
| v | |
| pretrained Transformer | |
| | | |
| shared representation / KV cache | |
| | | |
| +-------------------------------+ | |
| | | | | |
| question 1 question 2 question N | |
| | | | | |
| candidate set candidate set candidate set | |
| | | | | |
| v v v | |
| scalar logits scalar logits scalar logits | |
| | | | | |
| softmax / sigmoid / ordinal mapping | |
| | | |
| calibrated probabilities | |
| ``` | |
| The model should **not generate answer text**. | |
| It should directly compute probabilities over caller-provided decisions. | |
| Examples: | |
| ```text | |
| state: | |
| Customer has made 3 late payments... | |
| question: | |
| "Will this customer churn within 30 days?" | |
| output: | |
| p_yes = 0.71 | |
| ``` | |
| or: | |
| ```text | |
| question: | |
| "Which team should handle this ticket?" | |
| choices: | |
| - billing | |
| - security | |
| - infrastructure | |
| - sales | |
| output: | |
| billing 0.03 | |
| security 0.83 | |
| infrastructure 0.12 | |
| sales 0.02 | |
| ``` | |
| Choices must be dynamic natural-language values supplied at inference time. | |
| Do not implement the task as a fixed classifier head whose labels are predetermined at training time. | |
| --- | |
| # 3. Research questions | |
| The work should answer these questions. | |
| ## RQ1 — Can a normal pretrained decoder LLM provide the required capability? | |
| Start with an existing 1.5B–8B pretrained model rather than training a foundation model from scratch. | |
| Likely candidates: | |
| - Qwen-family base models; | |
| - Llama-family base models if practical; | |
| - Gemma-family base models; | |
| - another strong dense 3B–8B base model. | |
| Prefer a **base model rather than an instruction/chat model** for the main experiment if tooling permits. | |
| Test at least: | |
| - ~1.5B; | |
| - ~3B; | |
| - ~7B/8B. | |
| A 3B model is the most interesting target. | |
| ## RQ2 — Is explicit decision training necessary? | |
| Compare: | |
| A. unmodified model using candidate token logits; | |
| B. model fine-tuned for decision scoring; | |
| C. model with a dedicated scalar decision readout; | |
| D. optionally, model with a shallow decision decoder over shared state representations. | |
| ## RQ3 — Can the state computation be amortised? | |
| Measure latency as the number of questions grows. | |
| The desirable behaviour is roughly: | |
| ```text | |
| 1 question ~ X ms | |
| 10 questions ~ X + small overhead | |
| 100 questions ~ X + moderate overhead | |
| ``` | |
| rather than linear full-model execution per question. | |
| ## RQ4 — Can the probabilities be made genuinely calibrated? | |
| Do not equate softmax output with calibration. | |
| Evaluate: | |
| - expected calibration error; | |
| - Brier score; | |
| - log loss; | |
| - reliability diagrams; | |
| - selective accuracy / coverage; | |
| - calibration under domain shift. | |
| ## RQ5 — What is the best architectural split? | |
| Compare at least two architectures if feasible: | |
| ### Architecture A — Shared-prefix decoder | |
| Use the pretrained causal model normally for the state. | |
| Prefill once. | |
| Reuse the state KV cache across many independent branches. | |
| Each branch contains: | |
| ```text | |
| question + candidate | |
| ``` | |
| and yields one scalar score. | |
| ### Architecture B — Shared state encoder + shallow decision decoder | |
| Run the main Transformer once over state. | |
| Use the resulting state representation as memory. | |
| Run only a small number of layers for each question/candidate. | |
| Conceptually: | |
| ```text | |
| state | |
| | | |
| large Transformer | |
| | | |
| H_state | |
| | | |
| +--> short question/candidate decoder --> scalar | |
| +--> short question/candidate decoder --> scalar | |
| +--> ... | |
| ``` | |
| Architecture B is more speculative but potentially closer to the ideal latency characteristics. | |
| --- | |
| # 4. First implementation: baseline before training | |
| Build the cheapest useful reproduction first. | |
| Use an off-the-shelf small decoder model. | |
| ## Shared-prefix classifier | |
| For every request: | |
| 1. Tokenise `state`. | |
| 2. Prefill the model once. | |
| 3. Save KV cache. | |
| 4. For every `(question, candidate)` pair: | |
| - append a canonical decision prompt; | |
| - reuse the same state KV cache; | |
| - compute a score. | |
| 5. Normalise scores across candidates. | |
| Example canonical input: | |
| ```text | |
| STATE: | |
| {state} | |
| QUESTION: | |
| {question} | |
| CANDIDATE: | |
| {candidate} | |
| DECISION: | |
| ``` | |
| Initially score candidates using one of: | |
| - probability of a special positive token; | |
| - difference between positive and negative logits; | |
| - a small learned scalar head applied to the final hidden representation. | |
| The third is preferred. | |
| Implement batching so all candidate branches execute together. | |
| The purpose of this baseline is to establish: | |
| - achievable latency; | |
| - scaling behaviour; | |
| - memory footprint; | |
| - calibration before dedicated training. | |
| Do this before designing a complicated training stack. | |
| --- | |
| # 5. Preferred model interface | |
| Expose three primitives. | |
| ## Boolean / Noul-like decision | |
| Input: | |
| ```json | |
| { | |
| "state": "...", | |
| "question": "..." | |
| } | |
| ``` | |
| Output: | |
| ```json | |
| { | |
| "false": 0.21, | |
| "true": 0.79 | |
| } | |
| ``` | |
| ## Choice | |
| Input: | |
| ```json | |
| { | |
| "state": "...", | |
| "question": "...", | |
| "choices": [ | |
| "billing", | |
| "fraud", | |
| "security", | |
| "technical support" | |
| ] | |
| } | |
| ``` | |
| Output: | |
| ```json | |
| { | |
| "billing": 0.05, | |
| "fraud": 0.02, | |
| "security": 0.88, | |
| "technical support": 0.05 | |
| } | |
| ``` | |
| ## Ordinal score | |
| Input: | |
| ```json | |
| { | |
| "state": "...", | |
| "question": "How severe is the incident?", | |
| "levels": [ | |
| "none", | |
| "low", | |
| "medium", | |
| "high", | |
| "critical" | |
| ] | |
| } | |
| ``` | |
| Output: | |
| ```json | |
| { | |
| "distribution": [0.01, 0.04, 0.20, 0.50, 0.25], | |
| "expected_score": 3.94 | |
| } | |
| ``` | |
| Internally, all three should preferably reduce to the same primitive: | |
| ```text | |
| score(state, question, candidate) -> scalar logit | |
| ``` | |
| --- | |
| # 6. Candidate scoring architecture | |
| Preferred target: | |
| ```python | |
| logit = model.score( | |
| state=state, | |
| question=question, | |
| candidate=candidate | |
| ) | |
| ``` | |
| The candidate can be arbitrary natural language. | |
| Do not bind a candidate to a vocabulary token. | |
| A possible implementation: | |
| ```text | |
| Transformer final hidden state | |
| | | |
| candidate/question representation | |
| | | |
| projection / interaction | |
| | | |
| scalar | |
| ``` | |
| Possible scalar head: | |
| ```python | |
| Linear(hidden_size, 1) | |
| ``` | |
| applied to a special `<decision>` position. | |
| A more advanced version may use: | |
| ```text | |
| state representation | |
| | | |
| cross attention from candidate/question tokens | |
| | | |
| pooling | |
| | | |
| MLP | |
| | | |
| scalar | |
| ``` | |
| Keep the first implementation simple. | |
| --- | |
| # 7. Training data | |
| The model must learn the meta-task: | |
| > Given arbitrary context, arbitrary natural-language decision criteria, and arbitrary natural-language candidates, assign useful probabilities. | |
| This is substantially different from training a classifier on a fixed taxonomy. | |
| Create a heterogeneous synthetic and real dataset. | |
| Important categories: | |
| - sentiment; | |
| - topic classification; | |
| - NLI; | |
| - entailment; | |
| - routing; | |
| - moderation-like policy classification; | |
| - intent classification; | |
| - relevance judgement; | |
| - document matching; | |
| - fraud/risk-style decisions; | |
| - medical-style abstract reasoning datasets, excluding unsafe deployment claims; | |
| - legal issue classification; | |
| - code defect categorisation; | |
| - support ticket routing; | |
| - factual yes/no; | |
| - uncertain factual questions; | |
| - ranking/recommendation tasks; | |
| - ordinal severity; | |
| - forecasting datasets where outcomes are known. | |
| Use public datasets where licensing permits. | |
| Convert every source dataset into the universal decision schema. | |
| --- | |
| # 8. Synthetic transformation | |
| Convert ordinary examples such as: | |
| ```text | |
| text: "The package arrived damaged." | |
| label: complaint | |
| ``` | |
| into: | |
| ```text | |
| state: | |
| The package arrived damaged. | |
| question: | |
| Which type of customer message is this? | |
| choices: | |
| - complaint | |
| - sales enquiry | |
| - praise | |
| - account cancellation | |
| ``` | |
| Then aggressively vary: | |
| - wording of question; | |
| - wording of choices; | |
| - order of choices; | |
| - number of choices; | |
| - irrelevant distractors; | |
| - synonymous labels; | |
| - label descriptions instead of label names; | |
| - long and short contexts; | |
| - multiple questions sharing the same state. | |
| Candidate permutation is mandatory. | |
| Without permutation, the model may learn positional priors. | |
| --- | |
| # 9. Probability targets | |
| Whenever possible, train against a distribution rather than only one-hot labels. | |
| Sources include: | |
| - empirical frequencies; | |
| - multiple human annotations; | |
| - multiple teacher-model samples; | |
| - disagreement among teacher models; | |
| - synthetic uncertainty; | |
| - known probabilistic forecasting datasets. | |
| Example: | |
| ```text | |
| candidate A: 0.52 | |
| candidate B: 0.43 | |
| candidate C: 0.05 | |
| ``` | |
| Do not collapse this to: | |
| ```text | |
| A = 1 | |
| B = 0 | |
| C = 0 | |
| ``` | |
| unless only hard labels are available. | |
| --- | |
| # 10. Training losses | |
| Start with ordinary differentiable losses. | |
| For Boolean decisions: | |
| ```text | |
| binary cross entropy | |
| + | |
| optional Brier loss | |
| ``` | |
| For categorical choices: | |
| ```text | |
| cross entropy | |
| ``` | |
| or, with soft targets: | |
| ```text | |
| KL(target_distribution || predicted_distribution) | |
| ``` | |
| For ordinal scores: | |
| ```text | |
| distributional cross entropy | |
| + | |
| optional ordinal / earth mover distance loss | |
| ``` | |
| Consider adding calibration-aware terms only after establishing strong baseline accuracy. | |
| Do not begin with PPO, GRPO or another RL algorithm. | |
| The problem is directly differentiable. | |
| --- | |
| # 11. RLCD-like experiment | |
| TypeSafe has not disclosed the exact RLCD algorithm. | |
| Treat this only as an experiment. | |
| One plausible objective is a proper scoring rule based on observed outcomes. | |
| Candidates: | |
| ### Log score | |
| ```text | |
| reward = log p(actual_outcome) | |
| ``` | |
| ### Brier score | |
| ```text | |
| reward = -sum((p_i - y_i)^2) | |
| ``` | |
| These incentivise honest probability estimates in expectation. | |
| Compare: | |
| - plain cross entropy; | |
| - Brier-trained model; | |
| - mixed CE+Brier; | |
| - post-hoc temperature calibration; | |
| - optional RL optimisation against a proper scoring-rule reward. | |
| Only keep RL if it provides a measurable advantage. | |
| --- | |
| # 12. Teacher distillation | |
| A useful way to bootstrap decision intelligence is to generate a large synthetic dataset using stronger models. | |
| For each training example: | |
| ```text | |
| state | |
| question | |
| choices | |
| ``` | |
| ask one or more strong teacher models for probability estimates. | |
| Prefer collecting: | |
| ```text | |
| full probability distribution | |
| ``` | |
| instead of a single selected answer. | |
| Potential strategy: | |
| 1. ask several independent teachers; | |
| 2. collect multiple samples; | |
| 3. convert agreement/disagreement into a target distribution; | |
| 4. remove examples with obviously malformed outputs; | |
| 5. use real labelled outcomes where available to correct teacher bias. | |
| Do not assume teacher confidence is calibrated. | |
| Treat teacher probabilities as noisy soft supervision. | |
| --- | |
| # 13. Multi-question training | |
| This is important. | |
| Training examples should include: | |
| ```text | |
| one shared state | |
| + | |
| many independent questions | |
| ``` | |
| Example: | |
| ```text | |
| state = long customer conversation | |
| question 1: | |
| Will the customer churn? | |
| question 2: | |
| Should this be escalated? | |
| question 3: | |
| Which department should own this? | |
| question 4: | |
| How angry is the customer? | |
| question 5: | |
| Is fraud suspected? | |
| ``` | |
| The implementation should batch these branches while preventing branch-to-branch leakage. | |
| The model should behave approximately as though each question were evaluated independently against the same state. | |
| --- | |
| # 14. Branching attention | |
| Once the simple KV-cache version works, test a packed branching attention implementation. | |
| Conceptually: | |
| ```text | |
| [STATE] | |
| [QUESTION 1 + CANDIDATE A] | |
| [QUESTION 1 + CANDIDATE B] | |
| [QUESTION 1 + CANDIDATE C] | |
| [QUESTION 2 + CANDIDATE A] | |
| ... | |
| ``` | |
| Attention rules: | |
| - state tokens attend normally; | |
| - each branch can attend to the state; | |
| - each branch can attend to itself; | |
| - branches cannot attend to other branches. | |
| This should allow many decision branches inside one model invocation. | |
| Initially implement with an attention mask even if inefficient. | |
| Only write a custom kernel after profiling proves this is important. | |
| --- | |
| # 15. Benchmark suite | |
| Create a reproducible benchmark before serious optimisation. | |
| The benchmark should contain at least: | |
| ### Accuracy | |
| - binary classification; | |
| - 4-way classification; | |
| - 20-way classification; | |
| - 255-way classification; | |
| - ordinal scoring. | |
| ### Context sizes | |
| - 128 tokens; | |
| - 1k; | |
| - 4k; | |
| - 16k; | |
| - optionally 32k. | |
| ### Number of independent questions | |
| - 1; | |
| - 4; | |
| - 16; | |
| - 64; | |
| - 256. | |
| ### Candidates per question | |
| - 2; | |
| - 4; | |
| - 10; | |
| - 100; | |
| - 255. | |
| Measure: | |
| - first-request latency; | |
| - warm latency; | |
| - throughput; | |
| - GPU utilisation; | |
| - memory use; | |
| - state prefill time; | |
| - branch evaluation time; | |
| - probability quality. | |
| --- | |
| # 16. Baselines | |
| Compare the system against: | |
| ## Baseline 1 — ordinary generation | |
| Prompt the same base model: | |
| ```text | |
| Select one answer and return JSON. | |
| ``` | |
| Measure: | |
| - latency; | |
| - parse failures; | |
| - accuracy. | |
| ## Baseline 2 — constrained decoding | |
| Use grammar / JSON / token restrictions. | |
| ## Baseline 3 — token-logit classification | |
| Use the next-token probabilities of candidate labels. | |
| ## Baseline 4 — separate full forward pass per candidate | |
| This demonstrates the value of shared computation. | |
| ## Baseline 5 — dedicated decision scorer | |
| The proposed architecture. | |
| --- | |
| # 17. Calibration evaluation | |
| Calibration is a core part of this project. | |
| For binary tasks compute: | |
| - accuracy; | |
| - AUROC if relevant; | |
| - negative log likelihood; | |
| - Brier score; | |
| - expected calibration error; | |
| - maximum calibration error. | |
| Create reliability plots. | |
| Example bins: | |
| ```text | |
| predicted 0.0–0.1 -> actual frequency | |
| predicted 0.1–0.2 -> actual frequency | |
| ... | |
| ``` | |
| Test calibration separately on: | |
| - in-distribution examples; | |
| - unseen datasets; | |
| - paraphrased questions; | |
| - adversarial distractors; | |
| - long contexts; | |
| - low-information inputs. | |
| A model whose accuracy is good but whose confidence is systematically wrong should not be considered successful. | |
| --- | |
| # 18. Calibration methods | |
| Evaluate: | |
| - raw logits; | |
| - global temperature scaling; | |
| - per-task-family temperature scaling; | |
| - isotonic regression for evaluation purposes; | |
| - Platt scaling; | |
| - training with Brier loss; | |
| - entropy regularisation; | |
| - label smoothing. | |
| Prefer methods that generalise across unknown tasks. | |
| Avoid solutions that require task-specific calibration at deployment, because the intended product is zero-shot arbitrary decision-making. | |
| --- | |
| # 19. Hardware deployment strategy | |
| Recommended split: | |
| ## Spark A | |
| Main training / serving experiment. | |
| Keep the 3B–8B model entirely local where possible. | |
| ## Spark B | |
| Teacher inference / synthetic-data generation. | |
| ## Spark C | |
| Evaluation, alternate checkpoint training, or parallel data generation. | |
| For distributed training, test: | |
| - DDP; | |
| - FSDP where necessary; | |
| - DeepSpeed if useful. | |
| Do not introduce distributed training merely because three machines exist. | |
| A model that fits on one Spark should first be proven on one Spark. | |
| --- | |
| # 20. Precision | |
| Test: | |
| Training: | |
| - BF16; | |
| - LoRA; | |
| - QLoRA if necessary. | |
| Inference: | |
| - BF16 baseline; | |
| - FP8 if supported and accurate; | |
| - INT8; | |
| - 4-bit weight-only quantisation. | |
| Measure calibration before and after quantisation. | |
| A quantised classifier may preserve top-1 accuracy while damaging probability calibration. | |
| That distinction matters. | |
| --- | |
| # 21. Suggested development stages | |
| ## Stage 0 — Environment | |
| Verify: | |
| - CUDA; | |
| - PyTorch; | |
| - NCCL; | |
| - inter-node communication; | |
| - FlashAttention / compatible attention implementation; | |
| - Transformers; | |
| - vLLM or alternative only where useful. | |
| Record exact versions. | |
| ## Stage 1 — No-training proof of concept | |
| Implement shared-prefix candidate scoring. | |
| Goal: | |
| - one context; | |
| - many independent questions; | |
| - no generated text. | |
| Deliver metrics. | |
| ## Stage 2 — Scalar decision head | |
| Add a learned scalar readout. | |
| Fine-tune a small model on public classification datasets. | |
| Goal: | |
| - outperform token-logit baseline. | |
| ## Stage 3 — Universal decision training | |
| Create heterogeneous transformed dataset. | |
| Goal: | |
| - zero-shot generalisation to unseen task types. | |
| ## Stage 4 — Soft-target distillation | |
| Generate probabilistic teacher targets. | |
| Goal: | |
| - stronger judgement and uncertainty quality. | |
| ## Stage 5 — Calibration | |
| Run held-out calibration experiments. | |
| Goal: | |
| - substantially lower Brier/ECE than ordinary LLM confidence prompting. | |
| ## Stage 6 — Multi-question optimisation | |
| Implement branch batching / branch masks. | |
| Goal: | |
| - sublinear latency growth with question count. | |
| ## Stage 7 — 3B model | |
| Train the most promising design at ~3B. | |
| Goal: | |
| - practical single-Spark inference. | |
| ## Stage 8 — 7B/8B model | |
| Only proceed if 3B scaling results suggest it is worthwhile. | |
| ## Stage 9 — Kernel optimisation | |
| Profile first. | |
| Only build custom CUDA/CuTe kernels for verified bottlenecks. | |
| --- | |
| # 22. Initial target metrics | |
| These are engineering targets, not claims about TypeSafe. | |
| For a ~3B model on one Spark, aim initially for: | |
| ```text | |
| short context + 1 decision: | |
| <150 ms warm | |
| short context + 16 questions: | |
| <250 ms | |
| short context + 100 questions: | |
| <500 ms | |
| ``` | |
| These targets may need adjustment after profiling. | |
| More important than absolute latency is the scaling curve: | |
| ```text | |
| latency(100 questions) | |
| ``` | |
| should be dramatically less than: | |
| ```text | |
| 100 * latency(1 question) | |
| ``` | |
| --- | |
| # 23. Repository structure | |
| Suggested layout: | |
| ```text | |
| decision-model/ | |
| ├── README.md | |
| ├── pyproject.toml | |
| ├── configs/ | |
| ├── data/ | |
| │ ├── adapters/ | |
| │ ├── synthetic/ | |
| │ └── benchmarks/ | |
| ├── src/ | |
| │ ├── model/ | |
| │ │ ├── scorer.py | |
| │ │ ├── scalar_head.py | |
| │ │ ├── branching_attention.py | |
| │ │ └── calibration.py | |
| │ ├── inference/ | |
| │ │ ├── shared_prefix.py | |
| │ │ ├── batching.py | |
| │ │ └── server.py | |
| │ ├── training/ | |
| │ │ ├── datasets.py | |
| │ │ ├── losses.py | |
| │ │ ├── trainer.py | |
| │ │ └── distillation.py | |
| │ └── eval/ | |
| │ ├── accuracy.py | |
| │ ├── calibration.py | |
| │ ├── latency.py | |
| │ └── scaling.py | |
| ├── scripts/ | |
| ├── tests/ | |
| └── results/ | |
| ``` | |
| Every experiment must save: | |
| - config; | |
| - git commit; | |
| - model; | |
| - dataset version; | |
| - hardware; | |
| - software versions; | |
| - metrics; | |
| - raw benchmark results. | |
| --- | |
| # 24. API prototype | |
| Expose a minimal HTTP API. | |
| Example: | |
| ```http | |
| POST /decide | |
| ``` | |
| ```json | |
| { | |
| "state": "The server emitted 500 errors after the deployment...", | |
| "questions": [ | |
| { | |
| "type": "boolean", | |
| "question": "Should the deployment be rolled back?" | |
| }, | |
| { | |
| "type": "choice", | |
| "question": "What is the most likely cause?", | |
| "choices": [ | |
| "database migration", | |
| "network outage", | |
| "expired TLS certificate", | |
| "traffic spike" | |
| ] | |
| } | |
| ] | |
| } | |
| ``` | |
| Return: | |
| ```json | |
| { | |
| "results": [ | |
| { | |
| "probability": 0.84 | |
| }, | |
| { | |
| "distribution": { | |
| "database migration": 0.67, | |
| "network outage": 0.10, | |
| "expired TLS certificate": 0.08, | |
| "traffic spike": 0.15 | |
| } | |
| } | |
| ] | |
| } | |
| ``` | |
| Questions must not influence each other. | |
| --- | |
| # 25. Important failure modes | |
| Watch specifically for: | |
| ### Candidate-position bias | |
| The first or last choice gets systematically favoured. | |
| ### Label-token bias | |
| The model scores familiar words better merely because of token frequency. | |
| ### Length bias | |
| Longer candidate descriptions receive consistently different scores. | |
| ### Confidence collapse | |
| Predictions cluster around 0.5. | |
| ### Overconfidence | |
| Model assigns >0.95 far too frequently. | |
| ### Question leakage | |
| One question affects another question in the same batch. | |
| ### Distribution-shift failure | |
| Calibration works only on training datasets. | |
| ### Teacher imitation | |
| The student reproduces quirks of one teacher instead of learning a robust decision representation. | |
| ### Quantisation damage | |
| Accuracy remains stable but probabilities become badly calibrated. | |
| --- | |
| # 26. Things not to do initially | |
| Do not: | |
| - pretrain a Transformer from scratch; | |
| - build a diffusion LM; | |
| - write custom CUDA before obtaining profiles; | |
| - use an enormous >30B model; | |
| - perform multi-node tensor parallelism by default; | |
| - optimise only top-1 accuracy; | |
| - evaluate confidence using self-reported generated numbers; | |
| - assume raw softmax is calibrated; | |
| - assume TypeSafe's implementation uses any specific architecture. | |
| --- | |
| # 27. Most important experiments | |
| If time is limited, prioritise these five. | |
| ## Experiment 1 | |
| Qwen-like 1.5B model. | |
| Shared state KV cache. | |
| Candidate token logits. | |
| Measure latency scaling. | |
| ## Experiment 2 | |
| Same model with a learned scalar `<decision>` head. | |
| Fine-tune on several classification datasets. | |
| Compare accuracy and calibration. | |
| ## Experiment 3 | |
| Train using candidate descriptions instead of fixed labels. | |
| Evaluate zero-shot on unseen label sets. | |
| ## Experiment 4 | |
| Train on soft probability distributions and Brier/log-loss objectives. | |
| Compare calibration against ordinary instruction-prompt confidence. | |
| ## Experiment 5 | |
| Run 1, 4, 16, 64 and 256 questions over one context and measure whether computation amortises as expected. | |
| Those experiments will determine whether the project is worth scaling. | |
| --- | |
| # 28. Main architectural decision gate | |
| After the first experiments, explicitly decide between: | |
| ```text | |
| A. Shared KV decoder branch architecture | |
| ``` | |
| and | |
| ```text | |
| B. Shared state encoder + shallow cross-attention decision decoder | |
| ``` | |
| Choose based on measured: | |
| - latency; | |
| - accuracy; | |
| - training stability; | |
| - implementation complexity; | |
| - multi-question scaling. | |
| Do not choose based on similarity to TypeSafe marketing. | |
| --- | |
| # 29. Definition of success | |
| A convincing prototype should demonstrate all of: | |
| 1. Dynamic natural-language choices with no fixed label vocabulary. | |
| 2. Direct probability output without autoregressive answer generation. | |
| 3. Shared context computation across many questions. | |
| 4. Better latency scaling than independent LLM calls. | |
| 5. Calibration materially better than generated “confidence scores.” | |
| 6. Useful zero-shot performance on task families absent from training. | |
| 7. Inference of a 3B-class model entirely on one DGX Spark. | |
| 8. Reproducible benchmarks and ablations showing which architectural components matter. | |
| The ideal result is not necessarily a model matching Jev. | |
| A successful result would establish that: | |
| > a pretrained small language model can be transformed into a specialised probabilistic decision engine whose inference characteristics are qualitatively different from those of a normal generative LLM. | |
| --- | |
| # 30. First action | |
| Begin with the smallest falsifiable implementation. | |
| Use a ~1.5B pretrained model. | |
| Implement: | |
| ```text | |
| state prefill once | |
| + | |
| batched independent question/candidate branches | |
| + | |
| direct candidate logits | |
| ``` | |
| Create a benchmark with: | |
| ```text | |
| 1 state | |
| 1 / 4 / 16 / 64 questions | |
| 2 / 4 / 16 / 255 candidates | |
| ``` | |
| Record: | |
| ```text | |
| prefill latency | |
| branch latency | |
| total latency | |
| peak memory | |
| accuracy | |
| Brier score | |
| ECE | |
| ``` | |
| Only after those results exist should the model architecture or training pipeline become more complicated. | |