Decision-1.0-Lex-0.6B / SYSTEM_ONE.md
Xunzhuo's picture
Accelerate mixed SystemOne decisions with verified typed scheduling
ee8e74d verified
|
Raw History Blame
6.18 kB

One state. Many decisions.

Decision uses the System One API format: supply state, model and a map of typed questions; receive an answers map under the same question IDs. Noul returns the probability of yes, Choice selects from your named options, and Score evaluates an ordered rubric.

from decision_inference import SystemOne

# native is your loaded Decision checkpoint; see USAGE.md for loading.
client = SystemOne(native, batching="auto")
result = client.system_one(
    state={"message": "Please refund the duplicate charge. I need this fixed today."},
    questions={
        "refund_requested": {
            "type": "noul",
            "instructions": "Does the customer explicitly request a refund?",
        },
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this request?",
            "criteria": {"Billing": "Charges and refunds", "Support": "Technical problems"},
        },
        "urgency": {
            "type": "score",
            "instructions": "How urgent is the request?",
            "criteria": ["No deadline", "Needed soon", "Needed today"],
        },
    },
)
print(result["answers"])

client.evaluate(request) accepts the complete JSON request object, including model. A client is bound to its loaded checkpoint; a mismatched model is an error. Public Kai/Lex names are inferred from their verified native manifest. For your own fine-tune, use SystemOne(native, model="my-decision-model").

Batch the questions or the contexts

One request accepts up to 128 questions. For one set of questions across many independent contexts, use client.batch:

questions = {
    "refund_requested": {
        "type": "noul",
        "instructions": "Does the customer explicitly request a refund?",
    }
}
results = client.batch([
    {"model": client.model, "state": message, "questions": questions}
    for message in messages
])
# results[i] corresponds to messages[i]; question IDs can repeat across requests.

batch is a Decision Python extension around ordinary System One requests. It accepts up to 128 requests and 512 total decisions, with a combined 2 MiB input limit. Responses preserve request order and question order. The entire batch must pass validation and complete-input token admission before any forward. An overlength or malformed input fails the call without partial answers.

The default uses physical batches of up to eight. batching="auto" uses the released padding-aware scheduler: up to 32 consecutive same-type decisions when that adds no padding; otherwise it keeps B8. FP32 rounding can vary with physical batch shape. Each state/question pair still has its own encoder computation. One API call does not imply one forward or a shared state activation cache.

Typed fields

Type criteria Answer
noul Optional true / false descriptions type, noul
choice 2–255 named options; descriptions may be null type, choice, probabilities, confidence
score 2–10 ordered level descriptions type, score, legend, probabilities, confidence

State, instructions and descriptions accept strings, JSON objects or arrays. Structured values become deterministic compact JSON with sorted object keys; array order, Choice option order and Score level order are preserved. Question IDs are bookkeeping only. Choice names are part of the semantic input, including when their description is null. Score levels are indexed from zero; score preserves the native probability-weighted FP32 expectation. Structured Score descriptions appear as JSON strings in legend.

Every complete state/question/candidate sequence must fit 1,024 tokens. Nothing is truncated. The token count includes the repeated state for each question; usage.input_tokens sums these complete sequences and usage.output_tokens is zero because the model returns scores without generating text. These are local computation counts, not TypeSafe billing counts.

Decision's confidence is the largest candidate probability, matching its native runtime. It is not calibrated correctness. TypeSafe does not specify its own formula in the confidence documentation, so thresholds should not be transferred between models without validation. The shared request/answer schema does not claim identical weights, confidence values, hosted service limits or SDK behavior.

Fine-tune with the same inputs

Use the same conversion for training so structured inputs, option names and criteria have identical semantics at training and inference time:

import json
from pathlib import Path
from decision_finetune.system_one import system_one_training_rows

request = json.loads(Path("examples/system-one.json").read_text())
rows = system_one_training_rows(
    request,  # the same model/state/questions object used for inference
    targets={
        "refund_requested": {"probability": 1.0},
        "team": {"choice_id": "Billing"},
        "urgency": {"probabilities": [0.0, 0.0, 1.0]},
    },
    request_id="ticket-1001",
    source_id="support-tickets",
    component_id="customer-42",
)

Save the rows as JSONL for the existing fine-tuning CLI. Targets stay separate from state, instructions and criteria. Related examples share a source/component ID and stay in one split; the CLI checks component and exact input overlap. Optional hard_target_ids supplies separate evaluation labels under the same question IDs. Noul supports soft yes probabilities; Choice and Score support complete soft distributions in the original candidate order.

Default typed scheduling

The default SystemOne path groups complete, admitted questions by decision type in physical batches of eight, then restores the original request and question order. This works for many questions over one state and questions across multiple contexts. No API changes or application-side sorting are needed. The optional batching="auto" policy is unchanged. Measured mixed-question scaling.