opensysone / source /RESEARCH_BRIEF.md
andyshu's picture
Back up verified OpenSysOne training snapshot and pinned source
2d5c26a verified
|
Raw History Blame
24.9 kB

Handover: Build a Jev-like Decision Model on 3 DGX Sparks

Objective

Build and evaluate a small-scale reproduction of the core ideas behind TypeSafe.ai's Jev / β€œSystem One” model using approximately three NVIDIA DGX Sparks.

The objective is not to reproduce TypeSafe's proprietary implementation exactly. The objective is to determine whether the following design can be made practical:

A pretrained Transformer is converted from an autoregressive text generator into a low-latency, zero-shot decision model that evaluates arbitrary natural-language questions and candidate outcomes in parallel and returns calibrated probability distributions directly.

The primary success criterion is evidence that this architecture can outperform or materially improve upon ordinary prompt-and-generate classification in at least some combination of:

  • latency;
  • throughput;
  • probability calibration;
  • consistency;
  • cost per decision;
  • scaling to many independent questions over one shared context.

Treat architecture claims about Jev as hypotheses unless independently verified.


1. Available hardware

Assume access to:

  • 3 Γ— NVIDIA DGX Spark;
  • 128 GB unified memory per Spark;
  • GB10 GPU;
  • approximately 273 GB/s local memory bandwidth per node;
  • 200 Gbit/s ConnectX-7 networking between Sparks;
  • direct Spark-to-Spark topology if practical.

Do not assume that the combined 384 GB behaves as unified GPU memory.

Prefer models that fit entirely on one Spark for inference.

Use multiple Sparks primarily for:

  • distributed fine-tuning;
  • synthetic data generation;
  • evaluation;
  • teacher inference;
  • parallel experiments.

Avoid tensor-parallel inference across Sparks unless necessary because network bandwidth is much lower than local memory bandwidth.


2. Core hypothesis

The highest-priority architecture to test is:

shared state/context
       |
       v
pretrained Transformer
       |
 shared representation / KV cache
       |
       +-------------------------------+
       |               |               |
    question 1      question 2      question N
       |               |               |
 candidate set     candidate set     candidate set
       |               |               |
       v               v               v
 scalar logits     scalar logits     scalar logits
       |               |               |
   softmax / sigmoid / ordinal mapping
       |
 calibrated probabilities

The model should not generate answer text.

It should directly compute probabilities over caller-provided decisions.

Examples:

state:
    Customer has made 3 late payments...

question:
    "Will this customer churn within 30 days?"

output:
    p_yes = 0.71

or:

question:
    "Which team should handle this ticket?"

choices:
    - billing
    - security
    - infrastructure
    - sales

output:
    billing        0.03
    security       0.83
    infrastructure 0.12
    sales          0.02

Choices must be dynamic natural-language values supplied at inference time.

Do not implement the task as a fixed classifier head whose labels are predetermined at training time.


3. Research questions

The work should answer these questions.

RQ1 β€” Can a normal pretrained decoder LLM provide the required capability?

Start with an existing 1.5B–8B pretrained model rather than training a foundation model from scratch.

Likely candidates:

  • Qwen-family base models;
  • Llama-family base models if practical;
  • Gemma-family base models;
  • another strong dense 3B–8B base model.

Prefer a base model rather than an instruction/chat model for the main experiment if tooling permits.

Test at least:

  • ~1.5B;
  • ~3B;
  • ~7B/8B.

A 3B model is the most interesting target.

RQ2 β€” Is explicit decision training necessary?

Compare:

A. unmodified model using candidate token logits;

B. model fine-tuned for decision scoring;

C. model with a dedicated scalar decision readout;

D. optionally, model with a shallow decision decoder over shared state representations.

RQ3 β€” Can the state computation be amortised?

Measure latency as the number of questions grows.

The desirable behaviour is roughly:

1 question    ~ X ms
10 questions  ~ X + small overhead
100 questions ~ X + moderate overhead

rather than linear full-model execution per question.

RQ4 β€” Can the probabilities be made genuinely calibrated?

Do not equate softmax output with calibration.

Evaluate:

  • expected calibration error;
  • Brier score;
  • log loss;
  • reliability diagrams;
  • selective accuracy / coverage;
  • calibration under domain shift.

RQ5 β€” What is the best architectural split?

Compare at least two architectures if feasible:

Architecture A β€” Shared-prefix decoder

Use the pretrained causal model normally for the state.

Prefill once.

Reuse the state KV cache across many independent branches.

Each branch contains:

question + candidate

and yields one scalar score.

Architecture B β€” Shared state encoder + shallow decision decoder

Run the main Transformer once over state.

Use the resulting state representation as memory.

Run only a small number of layers for each question/candidate.

Conceptually:

state
  |
large Transformer
  |
H_state
  |
  +--> short question/candidate decoder --> scalar
  +--> short question/candidate decoder --> scalar
  +--> ...

Architecture B is more speculative but potentially closer to the ideal latency characteristics.


4. First implementation: baseline before training

Build the cheapest useful reproduction first.

Use an off-the-shelf small decoder model.

Shared-prefix classifier

For every request:

  1. Tokenise state.
  2. Prefill the model once.
  3. Save KV cache.
  4. For every (question, candidate) pair:
    • append a canonical decision prompt;
    • reuse the same state KV cache;
    • compute a score.
  5. Normalise scores across candidates.

Example canonical input:

STATE:
{state}

QUESTION:
{question}

CANDIDATE:
{candidate}

DECISION:

Initially score candidates using one of:

  • probability of a special positive token;
  • difference between positive and negative logits;
  • a small learned scalar head applied to the final hidden representation.

The third is preferred.

Implement batching so all candidate branches execute together.

The purpose of this baseline is to establish:

  • achievable latency;
  • scaling behaviour;
  • memory footprint;
  • calibration before dedicated training.

Do this before designing a complicated training stack.


5. Preferred model interface

Expose three primitives.

Boolean / Noul-like decision

Input:

{
  "state": "...",
  "question": "..."
}

Output:

{
  "false": 0.21,
  "true": 0.79
}

Choice

Input:

{
  "state": "...",
  "question": "...",
  "choices": [
    "billing",
    "fraud",
    "security",
    "technical support"
  ]
}

Output:

{
  "billing": 0.05,
  "fraud": 0.02,
  "security": 0.88,
  "technical support": 0.05
}

Ordinal score

Input:

{
  "state": "...",
  "question": "How severe is the incident?",
  "levels": [
    "none",
    "low",
    "medium",
    "high",
    "critical"
  ]
}

Output:

{
  "distribution": [0.01, 0.04, 0.20, 0.50, 0.25],
  "expected_score": 3.94
}

Internally, all three should preferably reduce to the same primitive:

score(state, question, candidate) -> scalar logit

6. Candidate scoring architecture

Preferred target:

logit = model.score(
    state=state,
    question=question,
    candidate=candidate
)

The candidate can be arbitrary natural language.

Do not bind a candidate to a vocabulary token.

A possible implementation:

Transformer final hidden state
        |
candidate/question representation
        |
projection / interaction
        |
scalar

Possible scalar head:

Linear(hidden_size, 1)

applied to a special <decision> position.

A more advanced version may use:

state representation
        |
cross attention from candidate/question tokens
        |
pooling
        |
MLP
        |
scalar

Keep the first implementation simple.


7. Training data

The model must learn the meta-task:

Given arbitrary context, arbitrary natural-language decision criteria, and arbitrary natural-language candidates, assign useful probabilities.

This is substantially different from training a classifier on a fixed taxonomy.

Create a heterogeneous synthetic and real dataset.

Important categories:

  • sentiment;
  • topic classification;
  • NLI;
  • entailment;
  • routing;
  • moderation-like policy classification;
  • intent classification;
  • relevance judgement;
  • document matching;
  • fraud/risk-style decisions;
  • medical-style abstract reasoning datasets, excluding unsafe deployment claims;
  • legal issue classification;
  • code defect categorisation;
  • support ticket routing;
  • factual yes/no;
  • uncertain factual questions;
  • ranking/recommendation tasks;
  • ordinal severity;
  • forecasting datasets where outcomes are known.

Use public datasets where licensing permits.

Convert every source dataset into the universal decision schema.


8. Synthetic transformation

Convert ordinary examples such as:

text: "The package arrived damaged."
label: complaint

into:

state:
    The package arrived damaged.

question:
    Which type of customer message is this?

choices:
    - complaint
    - sales enquiry
    - praise
    - account cancellation

Then aggressively vary:

  • wording of question;
  • wording of choices;
  • order of choices;
  • number of choices;
  • irrelevant distractors;
  • synonymous labels;
  • label descriptions instead of label names;
  • long and short contexts;
  • multiple questions sharing the same state.

Candidate permutation is mandatory.

Without permutation, the model may learn positional priors.


9. Probability targets

Whenever possible, train against a distribution rather than only one-hot labels.

Sources include:

  • empirical frequencies;
  • multiple human annotations;
  • multiple teacher-model samples;
  • disagreement among teacher models;
  • synthetic uncertainty;
  • known probabilistic forecasting datasets.

Example:

candidate A: 0.52
candidate B: 0.43
candidate C: 0.05

Do not collapse this to:

A = 1
B = 0
C = 0

unless only hard labels are available.


10. Training losses

Start with ordinary differentiable losses.

For Boolean decisions:

binary cross entropy
+
optional Brier loss

For categorical choices:

cross entropy

or, with soft targets:

KL(target_distribution || predicted_distribution)

For ordinal scores:

distributional cross entropy
+
optional ordinal / earth mover distance loss

Consider adding calibration-aware terms only after establishing strong baseline accuracy.

Do not begin with PPO, GRPO or another RL algorithm.

The problem is directly differentiable.


11. RLCD-like experiment

TypeSafe has not disclosed the exact RLCD algorithm.

Treat this only as an experiment.

One plausible objective is a proper scoring rule based on observed outcomes.

Candidates:

Log score

reward = log p(actual_outcome)

Brier score

reward = -sum((p_i - y_i)^2)

These incentivise honest probability estimates in expectation.

Compare:

  • plain cross entropy;
  • Brier-trained model;
  • mixed CE+Brier;
  • post-hoc temperature calibration;
  • optional RL optimisation against a proper scoring-rule reward.

Only keep RL if it provides a measurable advantage.


12. Teacher distillation

A useful way to bootstrap decision intelligence is to generate a large synthetic dataset using stronger models.

For each training example:

state
question
choices

ask one or more strong teacher models for probability estimates.

Prefer collecting:

full probability distribution

instead of a single selected answer.

Potential strategy:

  1. ask several independent teachers;
  2. collect multiple samples;
  3. convert agreement/disagreement into a target distribution;
  4. remove examples with obviously malformed outputs;
  5. use real labelled outcomes where available to correct teacher bias.

Do not assume teacher confidence is calibrated.

Treat teacher probabilities as noisy soft supervision.


13. Multi-question training

This is important.

Training examples should include:

one shared state
+
many independent questions

Example:

state = long customer conversation

question 1:
    Will the customer churn?

question 2:
    Should this be escalated?

question 3:
    Which department should own this?

question 4:
    How angry is the customer?

question 5:
    Is fraud suspected?

The implementation should batch these branches while preventing branch-to-branch leakage.

The model should behave approximately as though each question were evaluated independently against the same state.


14. Branching attention

Once the simple KV-cache version works, test a packed branching attention implementation.

Conceptually:

[STATE]

[QUESTION 1 + CANDIDATE A]
[QUESTION 1 + CANDIDATE B]
[QUESTION 1 + CANDIDATE C]

[QUESTION 2 + CANDIDATE A]
...

Attention rules:

  • state tokens attend normally;
  • each branch can attend to the state;
  • each branch can attend to itself;
  • branches cannot attend to other branches.

This should allow many decision branches inside one model invocation.

Initially implement with an attention mask even if inefficient.

Only write a custom kernel after profiling proves this is important.


15. Benchmark suite

Create a reproducible benchmark before serious optimisation.

The benchmark should contain at least:

Accuracy

  • binary classification;
  • 4-way classification;
  • 20-way classification;
  • 255-way classification;
  • ordinal scoring.

Context sizes

  • 128 tokens;
  • 1k;
  • 4k;
  • 16k;
  • optionally 32k.

Number of independent questions

  • 1;
  • 4;
  • 16;
  • 64;

Candidates per question

  • 2;
  • 4;
  • 10;
  • 100;

Measure:

  • first-request latency;
  • warm latency;
  • throughput;
  • GPU utilisation;
  • memory use;
  • state prefill time;
  • branch evaluation time;
  • probability quality.

16. Baselines

Compare the system against:

Baseline 1 β€” ordinary generation

Prompt the same base model:

Select one answer and return JSON.

Measure:

  • latency;
  • parse failures;
  • accuracy.

Baseline 2 β€” constrained decoding

Use grammar / JSON / token restrictions.

Baseline 3 β€” token-logit classification

Use the next-token probabilities of candidate labels.

Baseline 4 β€” separate full forward pass per candidate

This demonstrates the value of shared computation.

Baseline 5 β€” dedicated decision scorer

The proposed architecture.


17. Calibration evaluation

Calibration is a core part of this project.

For binary tasks compute:

  • accuracy;
  • AUROC if relevant;
  • negative log likelihood;
  • Brier score;
  • expected calibration error;
  • maximum calibration error.

Create reliability plots.

Example bins:

predicted 0.0–0.1 -> actual frequency
predicted 0.1–0.2 -> actual frequency
...

Test calibration separately on:

  • in-distribution examples;
  • unseen datasets;
  • paraphrased questions;
  • adversarial distractors;
  • long contexts;
  • low-information inputs.

A model whose accuracy is good but whose confidence is systematically wrong should not be considered successful.


18. Calibration methods

Evaluate:

  • raw logits;
  • global temperature scaling;
  • per-task-family temperature scaling;
  • isotonic regression for evaluation purposes;
  • Platt scaling;
  • training with Brier loss;
  • entropy regularisation;
  • label smoothing.

Prefer methods that generalise across unknown tasks.

Avoid solutions that require task-specific calibration at deployment, because the intended product is zero-shot arbitrary decision-making.


19. Hardware deployment strategy

Recommended split:

Spark A

Main training / serving experiment.

Keep the 3B–8B model entirely local where possible.

Spark B

Teacher inference / synthetic-data generation.

Spark C

Evaluation, alternate checkpoint training, or parallel data generation.

For distributed training, test:

  • DDP;
  • FSDP where necessary;
  • DeepSpeed if useful.

Do not introduce distributed training merely because three machines exist.

A model that fits on one Spark should first be proven on one Spark.


20. Precision

Test:

Training:

  • BF16;
  • LoRA;
  • QLoRA if necessary.

Inference:

  • BF16 baseline;
  • FP8 if supported and accurate;
  • INT8;
  • 4-bit weight-only quantisation.

Measure calibration before and after quantisation.

A quantised classifier may preserve top-1 accuracy while damaging probability calibration.

That distinction matters.


21. Suggested development stages

Stage 0 β€” Environment

Verify:

  • CUDA;
  • PyTorch;
  • NCCL;
  • inter-node communication;
  • FlashAttention / compatible attention implementation;
  • Transformers;
  • vLLM or alternative only where useful.

Record exact versions.

Stage 1 β€” No-training proof of concept

Implement shared-prefix candidate scoring.

Goal:

  • one context;
  • many independent questions;
  • no generated text.

Deliver metrics.

Stage 2 β€” Scalar decision head

Add a learned scalar readout.

Fine-tune a small model on public classification datasets.

Goal:

  • outperform token-logit baseline.

Stage 3 β€” Universal decision training

Create heterogeneous transformed dataset.

Goal:

  • zero-shot generalisation to unseen task types.

Stage 4 β€” Soft-target distillation

Generate probabilistic teacher targets.

Goal:

  • stronger judgement and uncertainty quality.

Stage 5 β€” Calibration

Run held-out calibration experiments.

Goal:

  • substantially lower Brier/ECE than ordinary LLM confidence prompting.

Stage 6 β€” Multi-question optimisation

Implement branch batching / branch masks.

Goal:

  • sublinear latency growth with question count.

Stage 7 β€” 3B model

Train the most promising design at ~3B.

Goal:

  • practical single-Spark inference.

Stage 8 β€” 7B/8B model

Only proceed if 3B scaling results suggest it is worthwhile.

Stage 9 β€” Kernel optimisation

Profile first.

Only build custom CUDA/CuTe kernels for verified bottlenecks.


22. Initial target metrics

These are engineering targets, not claims about TypeSafe.

For a ~3B model on one Spark, aim initially for:

short context + 1 decision:
    <150 ms warm

short context + 16 questions:
    <250 ms

short context + 100 questions:
    <500 ms

These targets may need adjustment after profiling.

More important than absolute latency is the scaling curve:

latency(100 questions)

should be dramatically less than:

100 * latency(1 question)

23. Repository structure

Suggested layout:

decision-model/
β”œβ”€β”€ README.md
β”œβ”€β”€ pyproject.toml
β”œβ”€β”€ configs/
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ adapters/
β”‚   β”œβ”€β”€ synthetic/
β”‚   └── benchmarks/
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ model/
β”‚   β”‚   β”œβ”€β”€ scorer.py
β”‚   β”‚   β”œβ”€β”€ scalar_head.py
β”‚   β”‚   β”œβ”€β”€ branching_attention.py
β”‚   β”‚   └── calibration.py
β”‚   β”œβ”€β”€ inference/
β”‚   β”‚   β”œβ”€β”€ shared_prefix.py
β”‚   β”‚   β”œβ”€β”€ batching.py
β”‚   β”‚   └── server.py
β”‚   β”œβ”€β”€ training/
β”‚   β”‚   β”œβ”€β”€ datasets.py
β”‚   β”‚   β”œβ”€β”€ losses.py
β”‚   β”‚   β”œβ”€β”€ trainer.py
β”‚   β”‚   └── distillation.py
β”‚   └── eval/
β”‚       β”œβ”€β”€ accuracy.py
β”‚       β”œβ”€β”€ calibration.py
β”‚       β”œβ”€β”€ latency.py
β”‚       └── scaling.py
β”œβ”€β”€ scripts/
β”œβ”€β”€ tests/
└── results/

Every experiment must save:

  • config;
  • git commit;
  • model;
  • dataset version;
  • hardware;
  • software versions;
  • metrics;
  • raw benchmark results.

24. API prototype

Expose a minimal HTTP API.

Example:

POST /decide
{
  "state": "The server emitted 500 errors after the deployment...",
  "questions": [
    {
      "type": "boolean",
      "question": "Should the deployment be rolled back?"
    },
    {
      "type": "choice",
      "question": "What is the most likely cause?",
      "choices": [
        "database migration",
        "network outage",
        "expired TLS certificate",
        "traffic spike"
      ]
    }
  ]
}

Return:

{
  "results": [
    {
      "probability": 0.84
    },
    {
      "distribution": {
        "database migration": 0.67,
        "network outage": 0.10,
        "expired TLS certificate": 0.08,
        "traffic spike": 0.15
      }
    }
  ]
}

Questions must not influence each other.


25. Important failure modes

Watch specifically for:

Candidate-position bias

The first or last choice gets systematically favoured.

Label-token bias

The model scores familiar words better merely because of token frequency.

Length bias

Longer candidate descriptions receive consistently different scores.

Confidence collapse

Predictions cluster around 0.5.

Overconfidence

Model assigns >0.95 far too frequently.

Question leakage

One question affects another question in the same batch.

Distribution-shift failure

Calibration works only on training datasets.

Teacher imitation

The student reproduces quirks of one teacher instead of learning a robust decision representation.

Quantisation damage

Accuracy remains stable but probabilities become badly calibrated.


26. Things not to do initially

Do not:

  • pretrain a Transformer from scratch;
  • build a diffusion LM;
  • write custom CUDA before obtaining profiles;
  • use an enormous >30B model;
  • perform multi-node tensor parallelism by default;
  • optimise only top-1 accuracy;
  • evaluate confidence using self-reported generated numbers;
  • assume raw softmax is calibrated;
  • assume TypeSafe's implementation uses any specific architecture.

27. Most important experiments

If time is limited, prioritise these five.

Experiment 1

Qwen-like 1.5B model.

Shared state KV cache.

Candidate token logits.

Measure latency scaling.

Experiment 2

Same model with a learned scalar <decision> head.

Fine-tune on several classification datasets.

Compare accuracy and calibration.

Experiment 3

Train using candidate descriptions instead of fixed labels.

Evaluate zero-shot on unseen label sets.

Experiment 4

Train on soft probability distributions and Brier/log-loss objectives.

Compare calibration against ordinary instruction-prompt confidence.

Experiment 5

Run 1, 4, 16, 64 and 256 questions over one context and measure whether computation amortises as expected.

Those experiments will determine whether the project is worth scaling.


28. Main architectural decision gate

After the first experiments, explicitly decide between:

A. Shared KV decoder branch architecture

and

B. Shared state encoder + shallow cross-attention decision decoder

Choose based on measured:

  • latency;
  • accuracy;
  • training stability;
  • implementation complexity;
  • multi-question scaling.

Do not choose based on similarity to TypeSafe marketing.


29. Definition of success

A convincing prototype should demonstrate all of:

  1. Dynamic natural-language choices with no fixed label vocabulary.
  2. Direct probability output without autoregressive answer generation.
  3. Shared context computation across many questions.
  4. Better latency scaling than independent LLM calls.
  5. Calibration materially better than generated β€œconfidence scores.”
  6. Useful zero-shot performance on task families absent from training.
  7. Inference of a 3B-class model entirely on one DGX Spark.
  8. Reproducible benchmarks and ablations showing which architectural components matter.

The ideal result is not necessarily a model matching Jev.

A successful result would establish that:

a pretrained small language model can be transformed into a specialised probabilistic decision engine whose inference characteristics are qualitatively different from those of a normal generative LLM.


30. First action

Begin with the smallest falsifiable implementation.

Use a ~1.5B pretrained model.

Implement:

state prefill once
        +
batched independent question/candidate branches
        +
direct candidate logits

Create a benchmark with:

1 state
1 / 4 / 16 / 64 questions
2 / 4 / 16 / 255 candidates

Record:

prefill latency
branch latency
total latency
peak memory
accuracy
Brier score
ECE

Only after those results exist should the model architecture or training pipeline become more complicated.