Instructions to use samatv256/mini-Jev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use samatv256/mini-Jev with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
mini-Jev Qwen3-0.6B - step 10,626
mini-Jev is a 0.6B agentic decision model specialized for finite-choice tool and action selection. Give it a state, a question, and candidate actions; it returns a probability for each candidate. It uses a Qwen3-0.6B backbone with an adapter and a compact decision head. Candidate scoring is permutation equivariant, and the model returns finite-choice probabilities rather than generated text or tool arguments.
This is an early 10,626-update checkpoint. Training is planned through 50,000 updates; more training and broader decision supervision are coming. This interim release is not the final checkpoint.
Release files and revisions
The current step-010626 checkpoint loads its adapter from adapter/, its decision-head weights from decision_head.safetensors, and its inference loader from inference.py. Use these with config.json and the separately downloaded Qwen3-0.6B base model.
The retained root model.safetensors and odm_mini.py belong to the previous v1 baseline. They are preserved only for history and backward reference; they are not the step 10,626 weights or loader. To use that baseline, pin the previous revision.
Measured results
| Evaluation | Result | Scope |
|---|---|---|
| BFCL V1 Multiple Function - Jev function-selection adaptation | 96.5% (193/200) | Selects the supplied function identity from 2-4 definitions. Not an official BFCL leaderboard score; does not generate arguments or execute calls. |
| Strict held-out Choice | 65.6% (63/96) | A selected, bucket-balanced validation slice, not the full validation split. |
| Bespoke OOD typed decisions | 72.0% (36/50) | Small synthetic sanity check with straightforward distractors; not a public benchmark. |
| LocalLLaMA/typed-decisions, full test set | 34.0% (680/2,000) | Choice 27.0%, Noul 52.2%, Score 25.6%. General typed decisions remain a weakness. |
These evaluations have different datasets and denominators; their percentages are not directly comparable. The BFCL adaptation uses the BFCL V1 Multiple Function AST examples. The typed-decision result uses the LocalLLaMA/typed-decisions test split.
Intended use and limits
The current strength is finite-choice tool and action selection from state and supplied candidates. The model is less reliable for general typed decisions and for large candidate sets. It is not a general probability or classification model, and its probabilities have not been calibrated for arbitrary domains. A selected 17+ candidate held-out slice scored 4/17, so large sets need particular care.
Use
A CUDA GPU with BF16 support is required for this 4-bit package. Install the tested dependencies from requirements.txt. The Qwen3-0.6B base model is downloaded separately on first load; this repository contains only the adapter and decision head.
from inference import MiniJev
model = MiniJev.load(".")
result = model.predict(
state={"user_goal": "Find tomorrow's weather forecast for Paris."},
question="Given the current state and available options,\nwhich option should be selected?",
question_type="choice",
answer_options=[
{"id": "weather", "type": "tool", "label": "get_weather_forecast",
"description": "Look up the weather forecast for a place and date."},
{"id": "invoice", "type": "tool", "label": "calculate_invoice_total",
"description": "Add amounts on an invoice."},
{"id": "email", "type": "tool", "label": "send_email",
"description": "Send an email message."},
],
)
print(result["selected_id"])
print(result["options"])
Option IDs are bookkeeping and are excluded from model text. The model scores each supplied option and normalizes probabilities over the finite set. Input branches are limited to 8,192 tokens.
Data attribution
Decision supervision for this checkpoint draws in part on Jev Decisions v1, a public dataset released under CC BY 4.0. Its public upstream datasets carry their own attribution and licensing terms; see the dataset card and upstream license notes.
License
Apache-2.0. The Qwen3-0.6B base model is separately licensed by Qwen under Apache-2.0. See LICENSE and the base model card.
- Downloads last month
- 248
Task type is invalid.