How to use from the
Use from the
Transformers library
# Gated model: Login with a HF token with gated access permission
hf auth login
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("zero-shot-classification", model="empiriolabsai/aplomb-1")
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM

processor = AutoProcessor.from_pretrained("empiriolabsai/aplomb-1")
model = AutoModelForMultimodalLM.from_pretrained("empiriolabsai/aplomb-1", device_map="auto")
Quick Links

Accept the EmpirioLabs Model License to access Aplomb 1

Free for research, education, evaluation, personal use, and internal commercial use by organizations with annual revenue under US$1 million. Offering the model as a hosted service, larger commercial use, and shipping it in products sold to others need a commercial license from EmpirioLabs.ai LLC (support@empiriolabs.ai). Distilling or training other models from it or its outputs is not permitted.

Log in or Sign Up to review the conditions and access this model content.

Aplomb 1

Model page  ·  Launch post  ·  API docs  ·  Playground

Aplomb 1

Aplomb 1 is our first model, a decision model. It reads text, JSON, images, video and audio, and answers typed questions about them with calibrated probabilities instead of generated text. Send a state of up to 1M tokens with up to 128 questions, and every answer comes back as a probability distribution with a confidence value, in one request. Aplomb 1 has 5.3 billion parameters.

Highlights

  • Every input type in one model: text, JSON objects and arrays, images, video and audio.
  • Up to 1M tokens: state plus question, in a single request.
  • Four question types: yes or no, a choice among 2 to 255 options, an ordered score of 2 to 10 levels, and tool selection (which function to call, with distributions over its enum and boolean arguments).
  • Not in the state: ask for it per question, and the answer includes the probability that the state does not contain what the question needs.
  • Calibrated: a stated 80% is right about 80% of the time (Figure 3).
  • Independent answers: questions never influence each other.
  • Zero data retention by default on the hosted API: EmpirioLabs does not retain the content of requests or responses.

Results

Measured by EmpirioLabs on 2026-09-30. Accuracy is in percent; every system we ran answered the same items. Bold marks the best score in a row.

Accuracy on decision benchmarks

Decision benchmarks (the Intern-Decision evaluation bundle)

Benchmark Aplomb 1 Jev (p) Intern-Decision-4B JevK5 (p) SemIf (p) Kev-4B OmniJev-4B
JevBench Easy 100.0 100.0 100.0 100.0 100.0 100.0 87.5
JevBench Original 95.8 98.6 100.0 97.2 98.6 93.1 72.2
JevBench Hard 77.5 72.1 71.2 73.9 61.3 54.1 52.3
Typed Decisions 74.0 73.3 80.8 64.5 62.8 67.0 61.3
ToolACE 89.4 91.3 96.5 81.0 85.2 87.4 88.1
AG News 88.8 89.6 90.8 89.1 89.2 89.7 89.6
WildJailbreak 95.6 96.3 90.2 90.5 92.5 93.5 92.3
Average 88.7 88.7 89.9 85.2 84.2 83.5 77.6

(p) As published by Intern-Decision for the same items; not run by us. Intern-Decision reports 90.02 for Intern-Decision-4B on this bundle; its column shows our run.

More decision benchmarks

Benchmark Aplomb 1 Intern-Decision-4B Kev-4B OmniJev-4B
Calibration pilot 70.8 58.3 69.8 53.1
ANLI R3 54.0 54.2 50.8 n/a
Banking77 (77 options) 74.6 n/a 84.0 68.4
CLINC150 (150 options) 87.0 n/a 77.8 66.8

Images (first 500 items of each set)

Benchmark Aplomb 1 Intern-Decision-4B OmniJev-4B
POPE 89.0 89.2 88.4
MMStar 68.6 65.2 64.4
AI2D 87.0 81.0 83.4
MMMU 57.2 57.0 56.0
HallusionBench 79.8 79.4 71.2

Video

Benchmark Aplomb 1 OmniJev-4B
TempCompass 79.3 73.6
Video-MME (short) 80.7 72.2
NExT-QA 82.6 79.8

Audio

Benchmark Aplomb 1
VocalSound (vocal sounds) 91.9
CREMA-D (emotion in speech) 78.7
ESC-50 (environmental sounds) 77.0
MMAU (audio questions) 68.3
MMSU (spoken language) 63.2

We added audio to Aplomb 1 ourselves; none of the other decision models in these tables takes audio. It recognizes vocal sounds, emotion in speech and everyday sounds, and answers questions about what is said. Each clip also gets a short note with the sounds heard in it and a transcript of any speech, written by two open models included in audio_notes/ (CED-base and Whisper large-v3-turbo). The results above include the notes, and the API writes them the same way.

The audio results were updated on October 6, 2026, with new audio weights and the audio notes; text, image and video results are unchanged. If you downloaded the model before then, download it again.

Languages (MASSIVE intent, zero-shot: the request is in one of 51 languages, the question and the 59 intent labels are in English, 100 test requests per language, one choice among all 59 intents)

Average 75.0 across 51 languages, English 91.0, 39 languages at or above 70. Per language: en 91 | nl 89 | fa 87 | zh-CN 87 | it 86 | pl 86 | ja 85 | ko 85 | nb 85 | de 84 | fr 83 | ms 83 | zh-TW 83 | es 82 | he 82 | hi 82 | sv 82 | da 81 | fi 81 | tr 81 | af 80 | el 80 | vi 80 | id 79 | pt 79 | ro 78 | hu 77 | az 76 | bn 76 | hy 75 | jv 75 | ru 75 | ur 75 | ar 74 | kn 72 | te 72 | lv 71 | sl 71 | ka 70 | th 69 | sq 68 | is 67 | tl 63 | ml 62 | mn 62 | am 59 | km 58 | cy 53 | my 51 | sw 49 | ta 44.

Intent accuracy in each of 51 languages

Long states (lookups in order ledgers of the given length)

State length (tokens) 16K 128K 240K 512K 1M
Aplomb 1 100.0 100.0 100.0 100.0 96.8

Accuracy as the input grows to 1M tokens

Stated confidence against observed accuracy

Decision Index 0.2.1 (our run of the official kit on 2026-10-05; not an entry on the published board)

Aplomb 1 scores 44.86 on Decision Index 0.2.1, which averages 38 public benchmarks in five areas, each chance-corrected (0 is random guessing, 100 is perfect). All 150,317 scoreable requests were answered. Area scores are on the same 0 to 100 scale.

As of October 6, 2026, Aplomb 1 is #1 among 4B models on the published board, and #1 among models up to 5.3B parameters on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH.

The 8 Decision Index benchmarks where Aplomb 1 has the top score among models up to 5.3B on the published board

Aplomb 1 JPT-4B
Decision Index 44.86 43.04
Knowledge & Reasoning 29.7 28.7
Language Understanding 48.3 52.5
Retrieval & Classification 49.5 45.0
Tools & Automation 62.4 57.2
Arts & Human Taste 33.9 25.8

JPT-4B is the best 4B model on the published board, data of 2026-09-28.

Bar chart of Decision Index 0.2.1 scores: Aplomb 1 44.86 on our run of the official kit, ahead of every 4B model on the published board; next are JPT-4B at 43.04 and Jet v6.2 at 42.60.

Disclosure: Aplomb 1's training data included the public train splits of WinoGrande and ContractNLI, two of the 38 benchmarks in the index. We could not fully verify that index items were excluded from the earliest training data.

How it compares

Aplomb 1 compared with other decision models on question types, tool calls with argument probabilities, the not-in-the-input answer, inputs, context window, 1M-token speed, API formats, price per 1M input tokens and open weights

What each decision model or API accepts and answers, from its own published documentation as of 2026-09-29, with prices as of 2026-10-06:

Aplomb 1 Jev 1.13 OpenAI Decisions API (preview) Intern-Decision-4B Laya
Question types Yes or no, choice, score, tool Yes or no, choice, score Choice among listed answers Yes or no, choice, score Yes or no, choice, score
Tool selection with argument probabilities Yes No Not documented No No
Not in the state answer Yes No Not documented No No
Inputs Text, JSON, images, video, audio Text Text, images Text, images Text
Window 1,000,000 tokens 64K tokens Not published 8,192 tokens 512 to 8,192 tokens
Price per 1M input tokens $0.02 $0.042 Not announced Self-hosted $0.02 on Runware
Open weights Yes No No Yes Yes

Quickstart

The hosted Aplomb 1 API runs on EmpirioLabs' own custom inference runtime for decision models: a short question takes about 15 ms of model time, a question with an image about 35 ms, and a 1M-token state about 3 seconds in fast long-context mode, which is available only on the EmpirioLabs API and in the Playground. The hosted API also accepts the OpenAI, Anthropic and Gemini request formats, so their SDKs can call Aplomb 1 directly (chat formats).

curl https://api.empiriolabs.ai/v1/decisions \
  -H "Authorization: Bearer $EMPIRIOLABS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "aplomb-1",
    "state": {"order": {"id": "B-44120", "status": "shipped", "payment": "unpaid"}},
    "questions": {
      "paid": {"type": "noul", "instructions": "Has order B-44120 been paid?", "abstain": true},
      "carrier": {"type": "choice", "instructions": "Which carrier is delivering order B-44120?",
                  "criteria": {"ups": null, "fedex": null, "dhl": null}, "abstain": true}
    }
  }'

paid comes back as a confident no. carrier comes back with a high abstain probability, because the record does not name a carrier. Images, video and audio go in the state as {"type": "audio", "url": "https://..."} or {"type": "audio", "data": "<base64>", "mime": "audio/wav"}. The full schema and limits are in the Decisions API guide.

Tool selection for agents

Give Aplomb 1 your agent's function tools and the current state, and it returns the tool to call next with a probability for every tool and a distribution for each enum and boolean argument. The agent acts when the probability is high and hands the step to a larger model when it is not. We found no other decision model that returns argument probabilities.

{"model": "aplomb-1",
 "state": {"ticket": "Order B-44120 arrived with a cracked screen. I want my money back."},
 "questions": {"next_step": {"type": "tool", "instructions": "Which tool should the support agent call next?",
   "tools": [
     {"type": "function", "function": {"name": "issue_refund", "parameters": {"type": "object", "properties": {
       "reason": {"type": "string", "enum": ["damaged", "late", "not_as_described"]},
       "full_refund": {"type": "boolean"}}}}},
     {"type": "function", "function": {"name": "track_package", "parameters": {"type": "object", "properties": {
       "order_id": {"type": "string"}}}}},
     {"type": "function", "function": {"name": "escalate_to_human", "parameters": {"type": "object", "properties": {}}}}]}}}

Answer from the hosted API (abridged, rounded):

{"tool": "issue_refund",
 "probabilities": {"issue_refund": 0.969, "none": 0.025, "escalate_to_human": 0.005, "track_package": 0.002},
 "arguments": {"issue_refund": {
   "reason": {"value": "damaged", "probabilities": {"damaged": 0.993, "not_as_described": 0.007, "late": 0.001}},
   "full_refund": {"value": true, "probability_true": 0.761}}},
 "open_arguments": {"track_package": ["order_id"]}}

Free-text arguments such as the order id are listed under open_arguments for the agent to fill.

Long states on the hosted API

The hosted API reads a state longer than 131,072 tokens in fast long-context mode by default: a 1M-token state takes about 3 seconds instead of about 111 seconds read in full, and on our long-context test sets (164 documents of 131K to 1M tokens) fast mode answered all 525 decisions correctly, against 95.2% for a full read. Send "long_context": "full" to read every token, for example when an answer depends on a count or on confirming that something never appears. A request that evaluates more than 8 questions is read in full. Fast long-context mode is available only on the EmpirioLabs API and in the Playground; the open weights and run_aplomb.py read every token. See the Decisions API guide.

Decisions correct by document length for fast mode and a full read

Seconds per decision by document length for fast mode and a full read

Run locally

pip install -U "transformers>=5.17" torch accelerate torchvision torchaudio torchcodec soundfile librosa flash-linear-attention
python run_aplomb.py --model empiriolabsai/aplomb-1 --request request.json --debias

request.json is a Decisions API request body. The weights are bf16 and need about 12 GB of accelerator memory. For audio, the script writes the same notes the API does with the models in audio_notes/ (about 2 GB more); --no-audio-notes answers from the audio alone. The hosted API runs on EmpirioLabs' own optimized inference runtime for decision models, so its answers can differ slightly from the script's.

Intended use

Classification, routing, moderation, extraction checks, tool selection and evaluation, anywhere a program needs a probability rather than prose. Aplomb 1 does not generate text or fill free-text tool arguments. Where a decision has legal or similarly significant effects on a person, keep a human in the loop.

Limits

  • State plus the longest question: up to 1,000,000 tokens.
  • Media per request: up to 16 images, 8 audio clips and 4 videos.

License

EmpirioLabs Model License 1.0. Research, education, evaluation, personal use and internal commercial use by organizations with annual revenue under US$1 million are free. Offering the model as a hosted service, commercial use above that threshold and distribution in products sold to others need a commercial license from support@empiriolabs.ai. Using the model or its outputs to train another model is not permitted. Built on Qwen/Qwen3.5-4B, with the audio encoder of Qwen/Qwen3-Omni-30B-A3B-Instruct (both Apache 2.0; see LICENSE-APACHE-2.0 and NOTICE). audio_notes/ holds CED-base (Apache 2.0), Whisper large-v3-turbo (MIT, LICENSE-MIT) and the AudioSet class labels (CC BY 4.0), unchanged; see NOTICE.

Downloads last month
54
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for empiriolabsai/aplomb-1

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(890)
this model