Image-Text-to-Text
Transformers
Safetensors
Korean
English
qwen3_5
qwen3.5
multimodal
vision-language
korean
cultural-heritage
ocr
tool-use
long-horizon
conversational
Instructions to use KETI-NLP/Qwen3.5-KETI-HAECHI-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KETI-NLP/Qwen3.5-KETI-HAECHI-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="KETI-NLP/Qwen3.5-KETI-HAECHI-27B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("KETI-NLP/Qwen3.5-KETI-HAECHI-27B") model = AutoModelForMultimodalLM.from_pretrained("KETI-NLP/Qwen3.5-KETI-HAECHI-27B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use KETI-NLP/Qwen3.5-KETI-HAECHI-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KETI-NLP/Qwen3.5-KETI-HAECHI-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KETI-NLP/Qwen3.5-KETI-HAECHI-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/KETI-NLP/Qwen3.5-KETI-HAECHI-27B
- SGLang
How to use KETI-NLP/Qwen3.5-KETI-HAECHI-27B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "KETI-NLP/Qwen3.5-KETI-HAECHI-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KETI-NLP/Qwen3.5-KETI-HAECHI-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "KETI-NLP/Qwen3.5-KETI-HAECHI-27B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KETI-NLP/Qwen3.5-KETI-HAECHI-27B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use KETI-NLP/Qwen3.5-KETI-HAECHI-27B with Docker Model Runner:
docker model run hf.co/KETI-NLP/Qwen3.5-KETI-HAECHI-27B
|
Download reports/BENCHMARK_GUIDE.md from KETI-NLP/Qwen3.5-KETI-HAECHI-27B: direct link, hf CLI and curl.
- Browser
- Download file 4.59 kB
-
https://huggingface.co/KETI-NLP/Qwen3.5-KETI-HAECHI-27B/resolve/main/reports/BENCHMARK_GUIDE.md
- Command line
-
hf download hf://KETI-NLP/Qwen3.5-KETI-HAECHI-27B/reports/BENCHMARK_GUIDE.md
-
curl -L -o BENCHMARK_GUIDE.md https://huggingface.co/KETI-NLP/Qwen3.5-KETI-HAECHI-27B/resolve/main/reports/BENCHMARK_GUIDE.md
4.59 kB
Benchmark Guide
Reading the numbers
All percentage deltas in the model report are Tuned minus that model's own official Base, in absolute percentage points. Qwen and HCX raw scores must not be mixed into a single architecture ranking because their official checkpoints, chat templates, vision policies and pretraining histories differ.
| Metric | Meaning | Direction | Important caveat |
|---|---|---|---|
| Exact / normalized exact | After evaluator normalization, prediction and canonical answer are identical | Higher | Correct content with extra prose can fail |
| Containment | Canonical target occurs inside normalized output | Higher | More permissive than exact; hallucinated extra claims may still pass |
| Format | Required answer wrapper/choice/short-answer form was parsed | Higher | Format success is not semantic correctness |
| Similarity | Evaluator's normalized string/semantic similarity in [0,1] | Higher | Near-name answers can score partially without being canonical |
| CER | Character edit distance divided by reference length | Lower | Can exceed 1 when insertion is heavy |
| Line recall | Fraction of reference OCR lines recovered | Higher | Does not penalize all extra text |
| Accuracy / aAcc | Correct rows divided by evaluated rows; Hallusion aAcc balances its paired structure | Higher | Compare only matched protocols |
| pass@1 | First generated program passes hidden tests | Higher | Sensitive to code extraction and runtime |
| ToolSandbox similarity | Milestone completion minus minefield penalties, aggregated by scenario | Higher | Continuous, not plain accuracy |
| Tau reward | End-to-end task reward from actions, state checks and communication | Higher | Domain/trial weighting must match |
Benchmark families
General vision-language
- MMBench DEV EN v1.1: English visual perception, relation, logic and knowledge multiple choice.
- MMStar / MMStar-KO: leakage-reduced multimodal core skills in English and Korean.
- KRETA: Korean-centric document, chart, scene and reasoning evaluation.
- MMMU-Pro 10c: college-level multidisciplinary visual problems with ten choices.
- HallusionBench: visual faithfulness and hallucination resistance; report
aAcc. - MathVista MINI: visual mathematical reasoning. It is available in the HCX comparison; the Qwen run retained predictions but not an authoritative scored row.
Mammoth OCR and heritage
- Heritage Multi asks attributes, Reverse selects the matching image, and Simple freely names a heritage object.
- OCR Font, Outdoor and Public cover rendered fonts, signs/scenes and administrative documents.
- Exact, containment, similarity, CER and line recall should be read together. Exact alone confounds recognition and output policy.
Cultural Focus and H400/HS100
- Cultural Focus identity is exact official-title recognition; view caption is strict caption/year matching.
overallaggregates both and is therefore much harder than identity alone. - H400 Direct is open-vocabulary held-out naming; H400 Hard is same-family multiple choice; Knowledge Image/Text test factual knowledge.
- HS100 Train measures memorization/fit, Unseen measures new-view transfer, and Simple is only the 102-row intersection with the original Simple style. These are not aliases for full Mammoth Heritage Simple (768 rows).
OpenCompass language
- Core: IFEval, AIME 2024/2025, PRM800K math, BBH, GPQA, MMLU-Pro, HumanEval, LiveCodeBench and long-context suites.
- Extra: ARC, BoolQ, COPA, CommonsenseQA, HellaSwag, OpenBookQA, PIQA, SIQA, WinoGrande, MMLU, GSM8K, MATH, MBPP, LAMBADA, DROP and Natural Questions.
- Korean: KMMLU, CSATQA, HAE-RAE, K2-Eval, KoBEST, KOBALT and Korean clinical QA.
- The first row in each score table is the configured aggregate, not another dataset.
Tool use and agents
- BFCL V4 parses function calls into ASTs and tests simple, multiple, parallel and multi-turn calls, including missing-function/parameter and long-context cases.
- ToolSandbox runs phone-like stateful scenarios with distractor tools, milestone goals and minefields.
- Tau2 evaluates airline, retail and telecom interactions. Tau3 adds revised environments and banking knowledge with repeated trials.
- BFCL's run directory retained score CSVs and diagnostic logs but not the raw generation corpus; paired output examples are therefore provided for Tau2/Tau3, while BFCL is documented with its category scores and observed empty-response diagnostics in the main report.