Jehosephat

Jehosephat is a non-autoregressive decision model that answers questions about text or structured data. It returns probability distributions over your supplied questions in one forward pass, without generating text. Several questions may be asked in parallel. Questions can select a category (choice), score ordered levels (score), or estimate the probability of a proposition (noul).

The model architcture is a Qwen3-1.7B-Base backbone with a rank-32 LoRA adapter and a shared scalar decision head.

Jehosephat is inspired by TypeSafe's Jev, and so the query schema draws on TypeSafe Primitives.

Prerequisites

pip install "torch==2.14.0" "transformers==5.17.0" "peft==0.21.0"

Use

For a walkthrough with runnable examples, open example.ipynb.

Load Jehosephat and ask named questions:

from transformers import AutoModel

model = AutoModel.from_pretrained(
    "rhizomatous/jehosephat",
    trust_remote_code=True,
)
answers = model.predict(
    state={"message": "My order arrived broken."},
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this message?",
            "criteria": {
                "returns": "Damaged products, refunds, and replacements",
                "shipping": "Delivery tracking and delays",
            },
        },
    },
)

The model weights download automatically on first use. Jehosephat uses an available NVIDIA or Apple GPU automatically, and runs on CPU otherwise.

state accepts text or JSON-serializable data. Answers use your question names; those names do not enter the model input. Every answer contains type and a probabilities map whose values sum to 1. You can ask different question types together:

answers = model.predict(
    state={"message": "My order arrived broken."},
    questions={
        "route": {
            "type": "choice",
            "instructions": "Which team should handle this message?",
            "criteria": {
                "returns": "Damaged products, refunds, and replacements",
                "shipping": "Delivery tracking and delays",
            },
        },
        "severity": {
            "type": "score",
            "instructions": "How serious is the customer's problem?",
            "criteria": [
                "No problem reported",
                "Minor inconvenience",
                "Product unusable",
            ],
        },
        "replace": {
            "type": "noul",
            "instructions": "Does the customer need a replacement?",
        },
    },
)
Type Criteria Answer
choice Map of option names to descriptions choice: the winning option name; confidence: its probability
score Ordered list of level descriptions score: expected zero-based level; confidence: the largest level probability; legend: level names mapped to descriptions
noul Optional descriptions under "false" and "true" noul: probability of true

Results are plain Python dictionaries.

Score's probability keys are "0", "1", and so on. Score includes a legend mapping its level numbers to your descriptions. A three-level Score ranges from 0 to 2 and may be fractional. Its confidence refers to the most likely level, not a probability that the numerical expected score is correct.

A request supports 1–32 questions, 2–255 options per question, and at most 5,376 packed tokens in total, including formatting, state, questions, and options.

Architecture

The shared state, question text, and option branches are packed into one sequence. Each option is scored using the shared state, its question’s instructions, and its own key and description. Other questions and options are excluded from its context using a causal tree attention mask.

One linear scalar head scores every option's decision marker. A softmax normalizes those scores within each question. Inference uses SDPA with a dense tree mask.

Training

Jehosephat was trained on the following annotated and generated data. Targets preserve human annotation distributions or exact probabilities calculated by enumeration. Training data is not included in this model distribution.

Source and attribution Adaptation used in training
GoEmotions, Demszky et al., 2020 Emotion Noul questions and eligible 28-option Choice questions; empirical annotation distributions
Civil Comments, Borkan et al., 2019 Seven toxicity-related Noul questions per comment; continuous annotation fractions
SNLI, Bowman et al., 2015 Premise-grouped entailment Choice questions; five-judgment distributions
DynaSent, Potts et al., 2021 Ordered sentiment Score questions; full vote distributions, excluding rows with mixed-sentiment votes
HelpSteer2, Wang et al., 2024 Five response-quality Score questions; individual human rating distributions
Procedural workflows and exact probability experiments, viv shaw, 2026 Generated shared-state questions with exact posterior or predictive targets

This data was used to train a LoRA. LoRA rank was 32, alpha 64, and dropout 0.05, applied to attention projections and MLP gate/up/down projections. Adapter learning rate was 0.0002; head learning rate was 0.001. Microbatches held two requests with four gradient-accumulation steps. Weight decay was 0.01 and warmup ratio 0.05. Training permuted questions and options and forbade state truncation. The loss was cross-entropy plus 0.25 times Brier loss, averaged over questions within each request and then over requests; ordinal RPS weight was zero.

Benchmarks

Classifier benchmark

Evaluated on the v1 and v2 suites of jabr/classifier-benchmark:

Suite Overall Choice Noul Score
v1 67.9% (53/78) 88.0% (22/25) 50.0% (13/26) 66.7% (18/27)
v2 62.6% (542/866) 77.3% (280/362) 50.2% (156/311) 54.9% (106/193)

Choice is graded by the winning option, Noul by a 0.5 probability threshold, and Score by the most probable level, rather than the fractional expected score.

JevBench public subset

Evaluated on the 231 public items from JevBench:

Tier Overall Choice Noul Score
Easy 97.9% (47/48) 100.0% (36/36) 91.7% (11/12) No cases
Standard 54.2% (39/72) 50.0% (18/36) 45.8% (11/24) 83.3% (10/12)
Hard 38.7% (43/111) 28.4% (19/67) 52.6% (20/38) 66.7% (4/6)
Overall 55.8% (129/231) 52.5% (73/139) 56.8% (42/74) 77.8% (14/18)

Limitations

Options are scored independently before normalization. An option such as “none of the above” cannot inspect which other alternatives were supplied. Overlapping options, comparisons between sibling options, and set-dependent complements can therefore fail. Use self-contained criteria and put information needed by every option in the shared state or question. Adding options changes normalization, but does not let existing branches reason about the new options.

Training primarily covers English and a limited set of generated and annotated tasks. The aggregate results are heavily influenced by Noul questions, which were dominant in the training data.

Development

From a source checkout, install dependencies and development tools with uv:

uv sync --locked

Pytest and Ruff are used for checks. The release-loading and attention isolation tests run the actual model on CPU and require enough memory for its weights and inference.

Attribution

Jehosephat is licensed under the Apache-2.0 license. See LICENCE.qwen3.md for the license of the underlying Qwen3 model.

Citation

If you use Jehosephat in your work, please cite:

@misc{shaw2026jehosephat,
  author = {shaw, viv},
  title = {Jehosephat},
  year = {2026},
  url = {https://huggingface.co/rhizomatous/jehosephat}
}
Downloads last month
34
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for rhizomatous/jehosephat

Adapter
(81)
this model