Instructions to use darrellbest/Hemmingway-1-VL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use darrellbest/Hemmingway-1-VL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="darrellbest/Hemmingway-1-VL") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("darrellbest/Hemmingway-1-VL") model = AutoModelForMultimodalLM.from_pretrained("darrellbest/Hemmingway-1-VL", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use darrellbest/Hemmingway-1-VL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "darrellbest/Hemmingway-1-VL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darrellbest/Hemmingway-1-VL", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/darrellbest/Hemmingway-1-VL
- SGLang
How to use darrellbest/Hemmingway-1-VL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "darrellbest/Hemmingway-1-VL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darrellbest/Hemmingway-1-VL", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "darrellbest/Hemmingway-1-VL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "darrellbest/Hemmingway-1-VL", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use darrellbest/Hemmingway-1-VL with Docker Model Runner:
docker model run hf.co/darrellbest/Hemmingway-1-VL
Download README.md from darrellbest/Hemmingway-1-VL: direct link, hf CLI and curl.
- Browser
- Download file 3.19 kB
-
https://huggingface.co/darrellbest/Hemmingway-1-VL/resolve/main/README.md
- Command line
-
hf download hf://darrellbest/Hemmingway-1-VL/README.md
-
curl -L -o README.md https://huggingface.co/darrellbest/Hemmingway-1-VL/resolve/main/README.md
license: cc-by-nc-4.0
base_model:
- Altworld/Hemmingway-1
base_model_relation: finetune
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
tags:
- qwen3.8
- hemmingway
- vision-language
- creative-writing
- tool-calling
Hemmingway-1 VL
Altworld/Hemmingway-1 (a Qwen3.8-27B fine-tune for everyday writing),
converted from its text-only Qwen3_5ForCausalLM layout back into the stock
Qwen/Qwen3.8-27B vision-language layout
(Qwen3_5ForConditionalGeneration), with Qwen3.8-27B's own vision tower restored. Every language-model weight is
Hemmingway-1's, unchanged. The layout is the one the Qwen3.8 tooling ecosystem expects (vLLM, llama.cpp mmproj,
Heretic, quantizers), and you can give it images again.
All credit for the model goes to Altworld. This is a format conversion, nothing more.
What changed, and proof that nothing else did
- Tensor names:
model.layers.*→model.language_model.layers.*(andembed_tokens,norm).lm_head.*and the multi-token-prediction headmtp.*are kept as-is. - Added: the 333
model.visual.*tensors from Qwen/Qwen3.8-27B, plus its image/video processor configs. - Config: Qwen3.8-27B's own
config.json. Hemmingway-1's text config is identical to itstext_config, field for field.
Verified:
| Check | Result |
|---|---|
| All 866 Hemmingway-1 tensors (incl. MTP), byte-compared | byte-identical, same dtypes |
| All 333 vision tensors vs Qwen3.8-27B | byte-identical |
| Logits, original vs this repo (transformers 5.17, bf16), on 5 prompts incl. a tool call and a tool result | bit-equal (max difference 0) |
| 48-token greedy continuations on the same prompts | identical, incl. the <tool_call> output |
The chat template (including tool calling) is Qwen3.8's, which Hemmingway-1 already uses unchanged.
Use
from transformers import AutoModelForImageTextToText, AutoTokenizer
tok = AutoTokenizer.from_pretrained("darrellbest/Hemmingway-1-VL")
model = AutoModelForImageTextToText.from_pretrained("darrellbest/Hemmingway-1-VL", dtype="auto", device_map="auto")
vllm serve darrellbest/Hemmingway-1-VL
Other formats, each tested: GGUF (BF16 / Q8_0 / Q6_K / Q4_K_M + mmproj), FP8 and NVFP4 for vLLM. Uncensored: darrellbest/Hemmingway-1-Heretic.
For fine-tuning, either layout works. Altworld's original text-only repo is the simplest for text/tool-call LoRA.
Since the language weights are identical, an adapter trained on one applies to the other once its module names get
the language_model. prefix.
Licence
Same as Hemmingway-1: CC BY-NC 4.0. Non-commercial use, with credit to Hemmingway-1 / Altworld. Commercial use needs an agreement with Altworld (luka@hemmingway.io). The vision tower comes from Qwen3.8-27B (Apache-2.0).