Download SYSTEM_ONE.md from vllm-sr/Decision-1.0-Lex-0.6B: direct link, hf CLI and curl.
- Browser
- Download file 6.18 kB
-
https://huggingface.co/vllm-sr/Decision-1.0-Lex-0.6B/resolve/e4a96e95ad08867f7114e1a3ed4d42dca4e2723e/SYSTEM_ONE.md
- Command line
-
hf download hf://vllm-sr/Decision-1.0-Lex-0.6B@e4a96e95ad08867f7114e1a3ed4d42dca4e2723e/SYSTEM_ONE.md
-
curl -L -o SYSTEM_ONE.md https://huggingface.co/vllm-sr/Decision-1.0-Lex-0.6B/resolve/e4a96e95ad08867f7114e1a3ed4d42dca4e2723e/SYSTEM_ONE.md
One state. Many decisions.
Decision uses the System One API format: supply
state, model and a map of typed questions; receive an answers map under
the same question IDs. Noul returns the probability of yes, Choice selects from
your named options, and Score evaluates an ordered rubric.
from decision_inference import SystemOne
# native is your loaded Decision checkpoint; see USAGE.md for loading.
client = SystemOne(native, batching="auto")
result = client.system_one(
state={"message": "Please refund the duplicate charge. I need this fixed today."},
questions={
"refund_requested": {
"type": "noul",
"instructions": "Does the customer explicitly request a refund?",
},
"team": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": {"Billing": "Charges and refunds", "Support": "Technical problems"},
},
"urgency": {
"type": "score",
"instructions": "How urgent is the request?",
"criteria": ["No deadline", "Needed soon", "Needed today"],
},
},
)
print(result["answers"])
client.evaluate(request) accepts the complete JSON request object, including
model. A client is bound to its loaded checkpoint; a mismatched model is an
error. Public Kai/Lex names are inferred from their verified native manifest.
For your own fine-tune, use SystemOne(native, model="my-decision-model").
Batch the questions or the contexts
One request accepts up to 128 questions. For one set of questions across many
independent contexts, use client.batch:
questions = {
"refund_requested": {
"type": "noul",
"instructions": "Does the customer explicitly request a refund?",
}
}
results = client.batch([
{"model": client.model, "state": message, "questions": questions}
for message in messages
])
# results[i] corresponds to messages[i]; question IDs can repeat across requests.
batch is a Decision Python extension around ordinary System One requests. It
accepts up to 128 requests and 512 total decisions, with a combined 2 MiB input
limit. Responses preserve request order and question order. The entire batch
must pass validation and complete-input token admission before any forward.
An overlength or malformed input fails the call without partial answers.
The default uses physical batches of up to eight. batching="auto" uses the
released padding-aware scheduler: up to 32 consecutive same-type decisions when
that adds no padding; otherwise it keeps B8. FP32 rounding can vary with physical
batch shape. Each state/question pair still has its own encoder computation.
One API call does not imply one forward or a shared state activation cache.
Typed fields
| Type | criteria |
Answer |
|---|---|---|
noul |
Optional true / false descriptions |
type, noul |
choice |
2–255 named options; descriptions may be null | type, choice, probabilities, confidence |
score |
2–10 ordered level descriptions | type, score, legend, probabilities, confidence |
State, instructions and descriptions accept strings, JSON objects or arrays.
Structured values become deterministic compact JSON with sorted object keys;
array order, Choice option order and Score level order are preserved. Question
IDs are bookkeeping only. Choice names are part of the semantic input, including
when their description is null. Score levels are indexed from zero; score
preserves the native probability-weighted FP32 expectation. Structured Score
descriptions appear as JSON strings in legend.
Every complete state/question/candidate sequence must fit 1,024 tokens.
Nothing is truncated. The token count includes the repeated state for each
question; usage.input_tokens sums these complete sequences and
usage.output_tokens is zero because the model returns scores without generating
text. These are local computation counts, not TypeSafe billing counts.
Decision's confidence is the largest candidate probability, matching its native
runtime. It is not calibrated correctness. TypeSafe does not specify its own
formula in the confidence documentation,
so thresholds should not be transferred between models without validation.
The shared request/answer schema does not claim identical weights, confidence
values, hosted service limits or SDK behavior.
Fine-tune with the same inputs
Use the same conversion for training so structured inputs, option names and criteria have identical semantics at training and inference time:
import json
from pathlib import Path
from decision_finetune.system_one import system_one_training_rows
request = json.loads(Path("examples/system-one.json").read_text())
rows = system_one_training_rows(
request, # the same model/state/questions object used for inference
targets={
"refund_requested": {"probability": 1.0},
"team": {"choice_id": "Billing"},
"urgency": {"probabilities": [0.0, 0.0, 1.0]},
},
request_id="ticket-1001",
source_id="support-tickets",
component_id="customer-42",
)
Save the rows as JSONL for the existing fine-tuning CLI. Targets
stay separate from state, instructions and criteria. Related examples share a
source/component ID and stay in one split; the CLI checks component and exact
input overlap. Optional hard_target_ids supplies separate evaluation labels
under the same question IDs. Noul supports soft yes probabilities; Choice and
Score support complete soft distributions in the original candidate order.
Default typed scheduling
The default SystemOne path groups complete, admitted questions by decision type in physical batches of eight, then restores the original request and question order. This works for many questions over one state and questions across multiple contexts. No API changes or application-side sorting are needed. The optional batching="auto" policy is unchanged. Measured mixed-question scaling.