Text Generation
Transformers
Safetensors
English
qwen3_5
image-text-to-text
qwen3.5
decision-model
classification
lora
together-ai
conversational
Instructions to use stay-mellow-ai/mev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use stay-mellow-ai/mev with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="stay-mellow-ai/mev") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("stay-mellow-ai/mev") model = AutoModelForMultimodalLM.from_pretrained("stay-mellow-ai/mev", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use stay-mellow-ai/mev with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "stay-mellow-ai/mev" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "stay-mellow-ai/mev", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/stay-mellow-ai/mev
- SGLang
How to use stay-mellow-ai/mev with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "stay-mellow-ai/mev" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "stay-mellow-ai/mev", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "stay-mellow-ai/mev" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "stay-mellow-ai/mev", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use stay-mellow-ai/mev with Docker Model Runner:
docker model run hf.co/stay-mellow-ai/mev
Add measured confidence
Browse files
README.md
CHANGED
|
@@ -126,9 +126,9 @@ APIs like [TypeSafe's](https://docs.typesafe.ai/api) ask typed `choice`, `noul`,
|
|
| 126 |
|
| 127 |
| Question type | mev options | Answer |
|
| 128 |
|---|---|---|
|
| 129 |
-
| `choice` with `criteria: {key: description}` | One option per criteria entry | Most likely option,
|
| 130 |
| `noul` (yes/no) | `A` = yes, `B` = no | Probability of `A` |
|
| 131 |
-
| `score` with ordered `criteria: [...]` | One option per level, in order | Probability-weighted mean of the level indexes |
|
| 132 |
|
| 133 |
```python
|
| 134 |
import math
|
|
@@ -153,13 +153,49 @@ def option_probabilities(model, tokenizer, prompt, options):
|
|
| 153 |
scores = {o["key"]: logits[tokenizer.convert_tokens_to_ids(o["label"])].item() for o in options}
|
| 154 |
total = sum(math.exp(s) for s in scores.values())
|
| 155 |
return {key: math.exp(s) / total for key, s in scores.items()}
|
|
|
|
|
|
|
|
|
|
| 156 |
```
|
| 157 |
|
| 158 |
Compared with a typed-question API:
|
| 159 |
|
| 160 |
- **One question per call.** mev was trained on single decisions. To ask several questions about the same state, make one call per question, in parallel if you like.
|
| 161 |
- **2 to 24 options.** That is the range seen in training. Larger option sets are untested.
|
| 162 |
-
- **
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 163 |
|
| 164 |
## Training
|
| 165 |
|
|
@@ -213,7 +249,7 @@ These are development sets that shaped the recipe, not an independent benchmark.
|
|
| 213 |
|
| 214 |
- Generic chat is not the intended interface and may produce prose.
|
| 215 |
- The model can be wrong. Do not use it as the only authority for high-impact decisions.
|
| 216 |
-
- Prompt injection, multilingual behavior,
|
| 217 |
|
| 218 |
## License
|
| 219 |
|
|
|
|
| 126 |
|
| 127 |
| Question type | mev options | Answer |
|
| 128 |
|---|---|---|
|
| 129 |
+
| `choice` with `criteria: {key: description}` | One option per criteria entry | Most likely option, a probability per option, and `confidence` |
|
| 130 |
| `noul` (yes/no) | `A` = yes, `B` = no | Probability of `A` |
|
| 131 |
+
| `score` with ordered `criteria: [...]` | One option per level, in order | Probability-weighted mean of the level indexes, and `confidence` |
|
| 132 |
|
| 133 |
```python
|
| 134 |
import math
|
|
|
|
| 153 |
scores = {o["key"]: logits[tokenizer.convert_tokens_to_ids(o["label"])].item() for o in options}
|
| 154 |
total = sum(math.exp(s) for s in scores.values())
|
| 155 |
return {key: math.exp(s) / total for key, s in scores.items()}
|
| 156 |
+
|
| 157 |
+
def confidence(probabilities):
|
| 158 |
+
return max(probabilities.values())
|
| 159 |
```
|
| 160 |
|
| 161 |
Compared with a typed-question API:
|
| 162 |
|
| 163 |
- **One question per call.** mev was trained on single decisions. To ask several questions about the same state, make one call per question, in parallel if you like.
|
| 164 |
- **2 to 24 options.** That is the range seen in training. Larger option sets are untested.
|
| 165 |
+
- **`confidence` is computed client-side.** The model doesn't return a confidence field. Take the probability of the chosen option; see [Confidence](#confidence) for how well it matches real accuracy.
|
| 166 |
+
|
| 167 |
+
## Confidence
|
| 168 |
+
|
| 169 |
+
mev's `confidence` is the probability of the chosen option, after normalizing over the task's option letters. On the main development set (1,000 decisions), it closely tracks real accuracy, with an expected calibration error of 0.038. It also separates right answers from wrong ones well, with an AUROC of 0.88.
|
| 170 |
+
|
| 171 |
+
Accuracy when acting only on answers above a confidence threshold:
|
| 172 |
+
|
| 173 |
+
| Confidence at or above | Share of decisions kept | Accuracy on those |
|
| 174 |
+
|---:|---:|---:|
|
| 175 |
+
| 0.50 | 92.7% | 89.0% |
|
| 176 |
+
| 0.70 | 77.1% | 96.2% |
|
| 177 |
+
| 0.80 | 69.3% | 97.4% |
|
| 178 |
+
| 0.90 | 60.5% | 98.0% |
|
| 179 |
+
| 0.95 | 50.6% | 98.6% |
|
| 180 |
+
| 0.99 | 31.3% | 100.0% |
|
| 181 |
+
|
| 182 |
+
Reliability by confidence band:
|
| 183 |
+
|
| 184 |
+
| Confidence band | Decisions | Mean confidence | Accuracy |
|
| 185 |
+
|---|---:|---:|---:|
|
| 186 |
+
| below 0.50 | 73 | 0.44 | 57.5% |
|
| 187 |
+
| 0.50 to 0.70 | 156 | 0.60 | 53.2% |
|
| 188 |
+
| 0.70 to 0.80 | 78 | 0.75 | 85.9% |
|
| 189 |
+
| 0.80 to 0.90 | 88 | 0.85 | 93.2% |
|
| 190 |
+
| 0.90 to 0.95 | 99 | 0.93 | 94.9% |
|
| 191 |
+
| 0.95 to 0.99 | 193 | 0.97 | 96.4% |
|
| 192 |
+
| 0.99 and above | 313 | 1.00 | 100.0% |
|
| 193 |
+
|
| 194 |
+
Between 0.7 and 0.95, mev is slightly underconfident: it is right more often than its confidence says. Between 0.5 and 0.7, it is overconfident. A common pattern is to act automatically above a threshold like 0.9 and send the rest to a fallback, such as a larger model or a person.
|
| 195 |
+
|
| 196 |
+
Other ways to compute confidence were measured too. The gap between the top two options ranks answers about as well (AUROC 0.88) but has a calibration error of 0.14. A score based on entropy, which measures how spread out the probabilities are, does worse on both (AUROC 0.85, calibration error 0.17).
|
| 197 |
+
|
| 198 |
+
These numbers come from development sets that shaped the training recipe. Calibration can shift on your own data, so check your threshold against a labeled sample before relying on it.
|
| 199 |
|
| 200 |
## Training
|
| 201 |
|
|
|
|
| 249 |
|
| 250 |
- Generic chat is not the intended interface and may produce prose.
|
| 251 |
- The model can be wrong. Do not use it as the only authority for high-impact decisions.
|
| 252 |
+
- Prompt injection, multilingual behavior, and out-of-distribution robustness have not been comprehensively evaluated. Calibration was measured only on the development sets.
|
| 253 |
|
| 254 |
## License
|
| 255 |
|