Instructions to use empiriolabsai/aplomb-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use empiriolabsai/aplomb-1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("zero-shot-classification", model="empiriolabsai/aplomb-1")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("empiriolabsai/aplomb-1") model = AutoModelForMultimodalLM.from_pretrained("empiriolabsai/aplomb-1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("zero-shot-classification", model="empiriolabsai/aplomb-1")# pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM
processor = AutoProcessor.from_pretrained("empiriolabsai/aplomb-1")
model = AutoModelForMultimodalLM.from_pretrained("empiriolabsai/aplomb-1", device_map="auto")Accept the EmpirioLabs Model License to access Aplomb 1
Free for research, education, evaluation, personal use, and internal commercial use by organizations with annual revenue under US$1 million. Offering the model as a hosted service, larger commercial use, and shipping it in products sold to others need a commercial license from EmpirioLabs.ai LLC (support@empiriolabs.ai). Distilling or training other models from it or its outputs is not permitted.
Log in or Sign Up to review the conditions and access this model content.
Model page · Launch post · API docs · Playground
Aplomb 1
Aplomb 1 is our first model, a decision model. It reads text, JSON, images, video and audio, and answers typed questions about them with calibrated probabilities instead of generated text. Send a state of up to 1M tokens with up to 128 questions, and every answer comes back as a probability distribution with a confidence value, in one request. Aplomb 1 has 5.3 billion parameters.
Highlights
- Every input type in one model: text, JSON objects and arrays, images, video and audio.
- Up to 1M tokens: state plus question, in a single request.
- Four question types: yes or no, a choice among 2 to 255 options, an ordered score of 2 to 10 levels, and tool
selection (which function to call, with distributions over its
enumandbooleanarguments). - Not in the state: ask for it per question, and the answer includes the probability that the state does not contain what the question needs.
- Calibrated: a stated 80% is right about 80% of the time (Figure 3).
- Independent answers: questions never influence each other.
- Zero data retention by default on the hosted API: EmpirioLabs does not retain the content of requests or responses.
Results
Measured by EmpirioLabs on 2026-09-30. Accuracy is in percent; every system we ran answered the same items. Bold marks the best score in a row.

Decision benchmarks (the Intern-Decision evaluation bundle)
| Benchmark | Aplomb 1 | Jev (p) | Intern-Decision-4B | JevK5 (p) | SemIf (p) | Kev-4B | OmniJev-4B |
|---|---|---|---|---|---|---|---|
| JevBench Easy | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 100.0 | 87.5 |
| JevBench Original | 95.8 | 98.6 | 100.0 | 97.2 | 98.6 | 93.1 | 72.2 |
| JevBench Hard | 77.5 | 72.1 | 71.2 | 73.9 | 61.3 | 54.1 | 52.3 |
| Typed Decisions | 74.0 | 73.3 | 80.8 | 64.5 | 62.8 | 67.0 | 61.3 |
| ToolACE | 89.4 | 91.3 | 96.5 | 81.0 | 85.2 | 87.4 | 88.1 |
| AG News | 88.8 | 89.6 | 90.8 | 89.1 | 89.2 | 89.7 | 89.6 |
| WildJailbreak | 95.6 | 96.3 | 90.2 | 90.5 | 92.5 | 93.5 | 92.3 |
| Average | 88.7 | 88.7 | 89.9 | 85.2 | 84.2 | 83.5 | 77.6 |
(p) As published by Intern-Decision for the same items; not run by us. Intern-Decision reports 90.02 for Intern-Decision-4B on this bundle; its column shows our run.
More decision benchmarks
| Benchmark | Aplomb 1 | Intern-Decision-4B | Kev-4B | OmniJev-4B |
|---|---|---|---|---|
| Calibration pilot | 70.8 | 58.3 | 69.8 | 53.1 |
| ANLI R3 | 54.0 | 54.2 | 50.8 | n/a |
| Banking77 (77 options) | 74.6 | n/a | 84.0 | 68.4 |
| CLINC150 (150 options) | 87.0 | n/a | 77.8 | 66.8 |
Images (first 500 items of each set)
| Benchmark | Aplomb 1 | Intern-Decision-4B | OmniJev-4B |
|---|---|---|---|
| POPE | 89.0 | 89.2 | 88.4 |
| MMStar | 68.6 | 65.2 | 64.4 |
| AI2D | 87.0 | 81.0 | 83.4 |
| MMMU | 57.2 | 57.0 | 56.0 |
| HallusionBench | 79.8 | 79.4 | 71.2 |
Video
| Benchmark | Aplomb 1 | OmniJev-4B |
|---|---|---|
| TempCompass | 79.3 | 73.6 |
| Video-MME (short) | 80.7 | 72.2 |
| NExT-QA | 82.6 | 79.8 |
Audio
| Benchmark | Aplomb 1 |
|---|---|
| VocalSound (vocal sounds) | 91.9 |
| CREMA-D (emotion in speech) | 78.7 |
| ESC-50 (environmental sounds) | 77.0 |
| MMAU (audio questions) | 68.3 |
| MMSU (spoken language) | 63.2 |
We added audio to Aplomb 1 ourselves; none of the other decision models in these tables takes audio. It recognizes
vocal sounds, emotion in speech and everyday sounds, and answers questions about what is said. Each clip also gets
a short note with the sounds heard in it and a transcript of any speech, written by two open models included in
audio_notes/ (CED-base and Whisper large-v3-turbo). The results above include the notes, and the API writes them
the same way.
The audio results were updated on October 6, 2026, with new audio weights and the audio notes; text, image and video results are unchanged. If you downloaded the model before then, download it again.
Languages (MASSIVE intent, zero-shot: the request is in one of 51 languages, the question and the 59 intent labels are in English, 100 test requests per language, one choice among all 59 intents)
Average 75.0 across 51 languages, English 91.0, 39 languages at or above 70. Per language: en 91 | nl 89 | fa 87 | zh-CN 87 | it 86 | pl 86 | ja 85 | ko 85 | nb 85 | de 84 | fr 83 | ms 83 | zh-TW 83 | es 82 | he 82 | hi 82 | sv 82 | da 81 | fi 81 | tr 81 | af 80 | el 80 | vi 80 | id 79 | pt 79 | ro 78 | hu 77 | az 76 | bn 76 | hy 75 | jv 75 | ru 75 | ur 75 | ar 74 | kn 72 | te 72 | lv 71 | sl 71 | ka 70 | th 69 | sq 68 | is 67 | tl 63 | ml 62 | mn 62 | am 59 | km 58 | cy 53 | my 51 | sw 49 | ta 44.

Long states (lookups in order ledgers of the given length)
| State length (tokens) | 16K | 128K | 240K | 512K | 1M |
|---|---|---|---|---|---|
| Aplomb 1 | 100.0 | 100.0 | 100.0 | 100.0 | 96.8 |


Decision Index 0.2.1 (our run of the official kit on 2026-10-05; not an entry on the published board)
Aplomb 1 scores 44.86 on Decision Index 0.2.1, which averages 38 public benchmarks in five areas, each chance-corrected (0 is random guessing, 100 is perfect). All 150,317 scoreable requests were answered. Area scores are on the same 0 to 100 scale.
As of October 6, 2026, Aplomb 1 is #1 among 4B models on the published board, and #1 among models up to 5.3B parameters on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH.

| Aplomb 1 | JPT-4B | |
|---|---|---|
| Decision Index | 44.86 | 43.04 |
| Knowledge & Reasoning | 29.7 | 28.7 |
| Language Understanding | 48.3 | 52.5 |
| Retrieval & Classification | 49.5 | 45.0 |
| Tools & Automation | 62.4 | 57.2 |
| Arts & Human Taste | 33.9 | 25.8 |
JPT-4B is the best 4B model on the published board, data of 2026-09-28.

Disclosure: Aplomb 1's training data included the public train splits of WinoGrande and ContractNLI, two of the 38 benchmarks in the index. We could not fully verify that index items were excluded from the earliest training data.
How it compares

What each decision model or API accepts and answers, from its own published documentation as of 2026-09-29, with prices as of 2026-10-06:
| Aplomb 1 | Jev 1.13 | OpenAI Decisions API (preview) | Intern-Decision-4B | Laya | |
|---|---|---|---|---|---|
| Question types | Yes or no, choice, score, tool | Yes or no, choice, score | Choice among listed answers | Yes or no, choice, score | Yes or no, choice, score |
| Tool selection with argument probabilities | Yes | No | Not documented | No | No |
| Not in the state answer | Yes | No | Not documented | No | No |
| Inputs | Text, JSON, images, video, audio | Text | Text, images | Text, images | Text |
| Window | 1,000,000 tokens | 64K tokens | Not published | 8,192 tokens | 512 to 8,192 tokens |
| Price per 1M input tokens | $0.02 | $0.042 | Not announced | Self-hosted | $0.02 on Runware |
| Open weights | Yes | No | No | Yes | Yes |
Quickstart
The hosted Aplomb 1 API runs on EmpirioLabs' own custom inference runtime for decision models: a short question takes about 15 ms of model time, a question with an image about 35 ms, and a 1M-token state about 3 seconds in fast long-context mode, which is available only on the EmpirioLabs API and in the Playground. The hosted API also accepts the OpenAI, Anthropic and Gemini request formats, so their SDKs can call Aplomb 1 directly (chat formats).
curl https://api.empiriolabs.ai/v1/decisions \
-H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "aplomb-1",
"state": {"order": {"id": "B-44120", "status": "shipped", "payment": "unpaid"}},
"questions": {
"paid": {"type": "noul", "instructions": "Has order B-44120 been paid?", "abstain": true},
"carrier": {"type": "choice", "instructions": "Which carrier is delivering order B-44120?",
"criteria": {"ups": null, "fedex": null, "dhl": null}, "abstain": true}
}
}'
paid comes back as a confident no. carrier comes back with a high abstain probability, because the record
does not name a carrier. Images, video and audio go in the state as {"type": "audio", "url": "https://..."} or
{"type": "audio", "data": "<base64>", "mime": "audio/wav"}. The full schema and limits are in the
Decisions API guide.
Tool selection for agents
Give Aplomb 1 your agent's function tools and the current state, and it returns the tool to call next with a
probability for every tool and a distribution for each enum and boolean argument. The agent acts when the
probability is high and hands the step to a larger model when it is not. We found no other decision model that
returns argument probabilities.
{"model": "aplomb-1",
"state": {"ticket": "Order B-44120 arrived with a cracked screen. I want my money back."},
"questions": {"next_step": {"type": "tool", "instructions": "Which tool should the support agent call next?",
"tools": [
{"type": "function", "function": {"name": "issue_refund", "parameters": {"type": "object", "properties": {
"reason": {"type": "string", "enum": ["damaged", "late", "not_as_described"]},
"full_refund": {"type": "boolean"}}}}},
{"type": "function", "function": {"name": "track_package", "parameters": {"type": "object", "properties": {
"order_id": {"type": "string"}}}}},
{"type": "function", "function": {"name": "escalate_to_human", "parameters": {"type": "object", "properties": {}}}}]}}}
Answer from the hosted API (abridged, rounded):
{"tool": "issue_refund",
"probabilities": {"issue_refund": 0.969, "none": 0.025, "escalate_to_human": 0.005, "track_package": 0.002},
"arguments": {"issue_refund": {
"reason": {"value": "damaged", "probabilities": {"damaged": 0.993, "not_as_described": 0.007, "late": 0.001}},
"full_refund": {"value": true, "probability_true": 0.761}}},
"open_arguments": {"track_package": ["order_id"]}}
Free-text arguments such as the order id are listed under open_arguments for the agent to fill.
Long states on the hosted API
The hosted API reads a state longer than 131,072 tokens in fast long-context mode by default: a 1M-token state takes about 3
seconds instead of about 111 seconds read in full, and on our long-context test sets
(164 documents of 131K to 1M tokens) fast mode answered all 525 decisions correctly, against 95.2% for a full read.
Send "long_context": "full" to read every token, for example when an answer depends on a count or on confirming
that something never appears. A request that evaluates more than 8 questions is read in full. Fast long-context mode
is available only on the EmpirioLabs API and in the Playground; the open weights and run_aplomb.py read every
token. See the Decisions API guide.


Run locally
pip install -U "transformers>=5.17" torch accelerate torchvision torchaudio torchcodec soundfile librosa flash-linear-attention
python run_aplomb.py --model empiriolabsai/aplomb-1 --request request.json --debias
request.json is a Decisions API request body. The weights are bf16 and need about 12 GB of accelerator memory.
For audio, the script writes the same notes the API does with the models in audio_notes/ (about 2 GB more);
--no-audio-notes answers from the audio alone.
The hosted API runs on EmpirioLabs' own optimized inference runtime for decision models, so its answers can differ
slightly from the script's.
Intended use
Classification, routing, moderation, extraction checks, tool selection and evaluation, anywhere a program needs a probability rather than prose. Aplomb 1 does not generate text or fill free-text tool arguments. Where a decision has legal or similarly significant effects on a person, keep a human in the loop.
Limits
- State plus the longest question: up to 1,000,000 tokens.
- Media per request: up to 16 images, 8 audio clips and 4 videos.
License
EmpirioLabs Model License 1.0. Research, education, evaluation, personal use and internal commercial use
by organizations with annual revenue under US$1 million are free. Offering the model as a hosted service, commercial
use above that threshold and distribution in products sold to others need a commercial license from
support@empiriolabs.ai. Using the model or its outputs to train another model is not permitted. Built on
Qwen/Qwen3.5-4B, with the audio encoder of Qwen/Qwen3-Omni-30B-A3B-Instruct (both Apache 2.0; see
LICENSE-APACHE-2.0 and NOTICE). audio_notes/ holds CED-base (Apache 2.0), Whisper large-v3-turbo (MIT,
LICENSE-MIT) and the AudioSet class labels (CC BY 4.0), unchanged; see NOTICE.
- Downloads last month
- 54
# Gated model: Login with a HF token with gated access permission hf auth login