How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="OmniJev/OneJev-27B-FP8")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
pipe(text=messages)
# Load model directly
from transformers import AutoProcessor, AutoModelForMultimodalLM

processor = AutoProcessor.from_pretrained("OmniJev/OneJev-27B-FP8")
model = AutoModelForMultimodalLM.from_pretrained("OmniJev/OneJev-27B-FP8", device_map="auto")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    },
]
inputs = processor.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
Quick Links

OneJev-27B-FP8, a Multimodal System One Decision Model

Hugging Face Demo GitHub Website Data License

OneJev-27B with its decoder weights in 8-bit floating point, one scale per weight row and per token: 30.4 GB instead of 54.7 GB, so it fits on one 48 GB GPU. On 229 test rows it gives the 16-bit model's answer on 226 (98.7%), accuracy 65.5 for 16-bit and 66.4 for 8-bit. It needs a GPU with FP8 (L40S, H100, H200 and newer).

On one H200 with a 1280x720 screenshot it answers 1 question in 168 ms and 10 questions in one request in 298 ms, against 189 ms and 324 ms for 16-bit.

Quick start

pip install "qev[torch] @ git+https://github.com/OmniJev/OneJev.git"
qev serve --model OmniJev/OneJev-27B-FP8
from qev import Client, Choice, Noul
from qev.media import data_uri

r = Client("http://localhost:8000").system_one(
    state={"task": "Pay the open invoice from ACME", "screen": "<image:1>"},
    media=[{"type": "image", "data": data_uri("screenshot.png")}],
    questions={"done": Noul("The invoice has been paid"),
               "next": Choice("What should the agent do next?", {"click": "click an element", "stop": "stop"})},
)

The server speaks TypeSafe's System One API plus a media field for images and video. More examples, the latency benchmark and the code are on GitHub.

All sizes

Model Base Weights
OneJev-0.8B Qwen3.5-0.8B 2.2 GB
OneJev-4B Qwen3.5-4B 10.4 GB
OneJev-9B Qwen3.5-9B 18.8 GB
OneJev-27B Qwen3.8-27B 54.7 GB
OneJev-27B-FP8 OneJev-27B in 8-bit 30.4 GB

All sizes are in the OneJev collection.

Citation

@misc{onejev2026,
  title        = {{OneJev}: A Multimodal System One Decision Model},
  author       = {{OmniJev Team}},
  year         = {2026},
  howpublished = {\url{https://github.com/OmniJev/OneJev}}
}

License

Apache 2.0

Downloads last month
30
Safetensors
Model size
27B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OmniJev/OneJev-27B-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(3)
this model

Dataset used to train OmniJev/OneJev-27B-FP8

Collection including OmniJev/OneJev-27B-FP8