Text Classification
Transformers
Safetensors
GGUF
English
qwen3_5
image-text-to-text
tinyjev
jev
decision-model
system-one
typed-decisions
ollama
qwen3.5
conversational
Instructions to use AnkitAI/TinyJev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AnkitAI/TinyJev-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="AnkitAI/TinyJev-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AnkitAI/TinyJev-4B") model = AutoModelForMultimodalLM.from_pretrained("AnkitAI/TinyJev-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AnkitAI/TinyJev-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AnkitAI/TinyJev-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf AnkitAI/TinyJev-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AnkitAI/TinyJev-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf AnkitAI/TinyJev-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AnkitAI/TinyJev-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AnkitAI/TinyJev-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AnkitAI/TinyJev-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AnkitAI/TinyJev-4B:Q4_K_M
Use Docker
docker model run hf.co/AnkitAI/TinyJev-4B:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use AnkitAI/TinyJev-4B with Ollama:
ollama run hf.co/AnkitAI/TinyJev-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use AnkitAI/TinyJev-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AnkitAI/TinyJev-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "AnkitAI/TinyJev-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use AnkitAI/TinyJev-4B with Docker Model Runner:
docker model run hf.co/AnkitAI/TinyJev-4B:Q4_K_M
- Lemonade
How to use AnkitAI/TinyJev-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AnkitAI/TinyJev-4B:Q4_K_M
Run and chat with the model
lemonade run user.TinyJev-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use AnkitAI/TinyJev-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AnkitAI/TinyJev-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default AnkitAI/TinyJev-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use AnkitAI/TinyJev-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf AnkitAI/TinyJev-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "AnkitAI/TinyJev-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
update readme
Browse files
README.md
CHANGED
|
@@ -24,7 +24,7 @@ datasets: [jaredpalmer/kev-suites]
|
|
| 24 |
<p>
|
| 25 |
<a href="https://github.com/ankit-aglawe/tinyjev">GitHub</a> ·
|
| 26 |
<a href="https://pypi.org/project/tinyjev/">PyPI</a> ·
|
| 27 |
-
<a href="https://huggingface.co/AnkitAI/
|
| 28 |
<a href="https://github.com/ankit-aglawe/tinyjev/tree/main/examples">Examples</a>
|
| 29 |
</p>
|
| 30 |
|
|
@@ -53,13 +53,13 @@ Latency is a base M1 (16 GB) via MLX, one forward pass per case.
|
|
| 53 |
|
| 54 |
| Model | Params | OD-500 | Gate 0.85 | ms / case | Weights |
|
| 55 |
|---|---:|---:|---|---:|---|
|
| 56 |
-
| <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev 0.6B** | 596M, 1.2 GB | 440 (88.0%) | 59% @ 98.0% | 85 | 🤗 [AnkitAI/
|
| 57 |
-
| <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev 4B** | 4.0B, 8.0 GB | 474 (94.8%) | 87% @ 99.1% | 628 | 🤗 [AnkitAI/
|
| 58 |
|
| 59 |
OD-500 is correct answers out of 500. Gate 0.85 is the share of decisions answered on its own at
|
| 60 |
confidence ≥ 0.85, and how often those were right. Calibration (ECE 0.071 vs 0.022), coverage at 2%
|
| 61 |
error (63% vs 92%) and transfer-v4 dev (0.625 vs 0.762) are on the benchmark page. Load either with
|
| 62 |
-
`tinyjev.load("
|
| 63 |
|
| 64 |
Both rows are fp16. Loading with `quantize=8` keeps the same weights in half the memory and changes
|
| 65 |
almost nothing: the 0.6B scores 440 at 90 ms, the 4B 473 at 845 ms, one answer in 500 different from
|
|
@@ -74,10 +74,10 @@ On OpenDecision's Original Choice 500, a suite of 25 domains that was not in the
|
|
| 74 |
| Model | Correct / 500 | Handled alone at confidence ≥ 0.85 |
|
| 75 |
|---|---:|---:|
|
| 76 |
| Claude Opus 5.5 (cloud, self-reported probabilities) | 496 | 477 at 100.0% |
|
| 77 |
-
| **
|
| 78 |
-
|
|
| 79 |
| Kev-0.8B (raw logits) | 463 | 186 at 100.0% |
|
| 80 |
-
|
|
| 81 |
|
| 82 |
353/375 on dev, 121/125 on holdout, 95% CI 0.928–0.966, ECE 0.022, Brier 0.071, coverage at 2% error 92%.
|
| 83 |
On Kev's transfer-v4 dev it scores 0.762 against 0.625 for the 0.6B (Kev-4B, a full fine-tune, 0.790).
|
|
@@ -97,7 +97,7 @@ pip install 'tinyjev[torch]' # everything else
|
|
| 97 |
|
| 98 |
```python
|
| 99 |
import tinyjev
|
| 100 |
-
agent = tinyjev.load("
|
| 101 |
|
| 102 |
agent.predict({
|
| 103 |
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
|
|
@@ -116,12 +116,12 @@ Eight bits changed one answer in 500 on the held-out suite. Serve it over HTTP,
|
|
| 116 |
One request shape:
|
| 117 |
|
| 118 |
```bash
|
| 119 |
-
tinyjev serve
|
| 120 |
```
|
| 121 |
|
| 122 |
## What is in this repo
|
| 123 |
|
| 124 |
-
`AutoModel.from_pretrained("AnkitAI/
|
| 125 |
`Qwen3Model` in fp16 with the LoRA already merged. The decision head lives in `head.safetensors`, and
|
| 126 |
`tinyjev` is what turns hidden states into calibrated answers.
|
| 127 |
|
|
@@ -134,8 +134,8 @@ training. Temperature 1.0 at inference; the calibration figures above are at raw
|
|
| 134 |
|
| 135 |
| | transfer-v4 dev | decision-v7 dev | OpenDecision 500 |
|
| 136 |
| --- | --- | --- | --- |
|
| 137 |
-
|
|
| 138 |
-
|
|
| 139 |
|
| 140 |
## Support the Project
|
| 141 |
|
|
|
|
| 24 |
<p>
|
| 25 |
<a href="https://github.com/ankit-aglawe/tinyjev">GitHub</a> ·
|
| 26 |
<a href="https://pypi.org/project/tinyjev/">PyPI</a> ·
|
| 27 |
+
<a href="https://huggingface.co/AnkitAI/TinyJev-0.6B">TinyJev 0.6B</a> ·
|
| 28 |
<a href="https://github.com/ankit-aglawe/tinyjev/tree/main/examples">Examples</a>
|
| 29 |
</p>
|
| 30 |
|
|
|
|
| 53 |
|
| 54 |
| Model | Params | OD-500 | Gate 0.85 | ms / case | Weights |
|
| 55 |
|---|---:|---:|---|---:|---|
|
| 56 |
+
| <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev 0.6B** | 596M, 1.2 GB | 440 (88.0%) | 59% @ 98.0% | 85 | 🤗 [AnkitAI/TinyJev-0.6B](https://huggingface.co/AnkitAI/TinyJev-0.6B) |
|
| 57 |
+
| <img src="https://raw.githubusercontent.com/ankit-aglawe/tinyjev/main/assets/logos/tinyjev.png" width="18"> **TinyJev 4B** | 4.0B, 8.0 GB | 474 (94.8%) | 87% @ 99.1% | 628 | 🤗 [AnkitAI/TinyJev-4B](https://huggingface.co/AnkitAI/TinyJev-4B) |
|
| 58 |
|
| 59 |
OD-500 is correct answers out of 500. Gate 0.85 is the share of decisions answered on its own at
|
| 60 |
confidence ≥ 0.85, and how often those were right. Calibration (ECE 0.071 vs 0.022), coverage at 2%
|
| 61 |
error (63% vs 92%) and transfer-v4 dev (0.625 vs 0.762) are on the benchmark page. Load either with
|
| 62 |
+
`tinyjev.load("TinyJev-0.6B")` or `tinyjev.load("TinyJev-4B")`.
|
| 63 |
|
| 64 |
Both rows are fp16. Loading with `quantize=8` keeps the same weights in half the memory and changes
|
| 65 |
almost nothing: the 0.6B scores 440 at 90 ms, the 4B 473 at 845 ms, one answer in 500 different from
|
|
|
|
| 74 |
| Model | Correct / 500 | Handled alone at confidence ≥ 0.85 |
|
| 75 |
|---|---:|---:|
|
| 76 |
| Claude Opus 5.5 (cloud, self-reported probabilities) | 496 | 477 at 100.0% |
|
| 77 |
+
| **TinyJev-4B** (fp16) | **474** | **437 at 99.1%** |
|
| 78 |
+
| TinyJev-4B at `quantize=8` | 473 | 437 at 99.1% |
|
| 79 |
| Kev-0.8B (raw logits) | 463 | 186 at 100.0% |
|
| 80 |
+
| TinyJev-0.6B | 440 | 296 at 98.0% |
|
| 81 |
|
| 82 |
353/375 on dev, 121/125 on holdout, 95% CI 0.928–0.966, ECE 0.022, Brier 0.071, coverage at 2% error 92%.
|
| 83 |
On Kev's transfer-v4 dev it scores 0.762 against 0.625 for the 0.6B (Kev-4B, a full fine-tune, 0.790).
|
|
|
|
| 97 |
|
| 98 |
```python
|
| 99 |
import tinyjev
|
| 100 |
+
agent = tinyjev.load("TinyJev-4B", quantize=8) # 4.5 GB in memory; drop quantize for fp16, 8 GB
|
| 101 |
|
| 102 |
agent.predict({
|
| 103 |
"state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges on my card.",
|
|
|
|
| 116 |
One request shape:
|
| 117 |
|
| 118 |
```bash
|
| 119 |
+
tinyjev serve TinyJev-4B --quantize 8 # POST /v1/systemone on 127.0.0.1:8077
|
| 120 |
```
|
| 121 |
|
| 122 |
## What is in this repo
|
| 123 |
|
| 124 |
+
`AutoModel.from_pretrained("AnkitAI/TinyJev-4B")` loads the backbone on its own, a standard
|
| 125 |
`Qwen3Model` in fp16 with the LoRA already merged. The decision head lives in `head.safetensors`, and
|
| 126 |
`tinyjev` is what turns hidden states into calibrated answers.
|
| 127 |
|
|
|
|
| 134 |
|
| 135 |
| | transfer-v4 dev | decision-v7 dev | OpenDecision 500 |
|
| 136 |
| --- | --- | --- | --- |
|
| 137 |
+
| TinyJev-4B | 0.762 | 0.859 | 474 / 500 |
|
| 138 |
+
| TinyJev-0.6B | 0.625 | | 440 / 500 |
|
| 139 |
|
| 140 |
## Support the Project
|
| 141 |
|