Instructions to use internlm/Intern-Decision-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use internlm/Intern-Decision-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="internlm/Intern-Decision-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("internlm/Intern-Decision-4B") model = AutoModelForMultimodalLM.from_pretrained("internlm/Intern-Decision-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use internlm/Intern-Decision-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "internlm/Intern-Decision-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "internlm/Intern-Decision-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/internlm/Intern-Decision-4B
- SGLang
How to use internlm/Intern-Decision-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "internlm/Intern-Decision-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "internlm/Intern-Decision-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "internlm/Intern-Decision-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "internlm/Intern-Decision-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use internlm/Intern-Decision-4B with Docker Model Runner:
docker model run hf.co/internlm/Intern-Decision-4B
Not bad, but still not Jev.
Greetings,
Running a test on human-calibrated data (my own email, souped and specific decision-relevant fields extracted) this doesn't match up to Jev. It does the same thing as a base Qwen 3.5-4B does (my own first test comparison as well): it over marks emails as 'Important', especially for emphatic lying, e.g. "Important news about your account, Roger" which is an AT&T marketing email.
Same input, same labels, lower accuracy on both my triage sets, half the 'Important' precision, worse calibration, and twice the false positives. At first I thought it was because I'd tuned the length of what I was sending to get better accuracy from Jev, so I shortened it for this model, and it got much worse. I'm running some more tests, this time with few-shot examples. (Fwiw, Jev got much worse with few-shot examples, focusing on the examples, not the 'things that are considered important' prompt. But an LLM-based model may behave differently. I'll comment if it changes anything.)
Not dismissing the project, just sharing my own testing results. I'd love to find something that competes with Jev, to run locally.
Good luck!
Thanks for your feedback!
In the email classification scenario, it seems the model follows some superficial patterns from the text itself (like Important news). We actually didn't mix too much classic classification task data into our training pipeline, and I think that's why it underperforms in your test.
This is a good chance for us to improve Intern-Decision. We'll collect relevant dataset to train a new version of the model. Hope future version would help!