Image-Text-to-Text
Transformers
Safetensors
English
glm5_next
shapleymcg
glm
exl3
tr3
vllm
quantized
dgx-spark
glm-5.3-flash
tensor-parallel
expert-parallel
speculative-decoding
serving-recipe
reproducibility
conversational
4-bit precision
Instructions to use jakejharris/jspark3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jakejharris/jspark3 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="jakejharris/jspark3") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("jakejharris/jspark3") model = AutoModelForMultimodalLM.from_pretrained("jakejharris/jspark3", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jakejharris/jspark3 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jakejharris/jspark3" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jakejharris/jspark3", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/jakejharris/jspark3
- SGLang
How to use jakejharris/jspark3 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jakejharris/jspark3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jakejharris/jspark3", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jakejharris/jspark3" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jakejharris/jspark3", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use jakejharris/jspark3 with Docker Model Runner:
docker model run hf.co/jakejharris/jspark3
Third-party acknowledgements
This repository is an independent research implementation. It does not wholesale-vendor the projects below, but its design and intended experiments build on their work. The BTX writer is the disclosed derived exception.
- Robert J. Aumann and Lloyd S. Shapley for the Aumann-Shapley value.
- Joshua Hill and NVIDIA Model Optimizer PR #2183 for Aumann-Shapley quantization sensitivity and the associated coverage/additivity analysis.
- Albert Tseng, Qingyao Sun, David Hou, Christopher De Sa, and the QTIP/QuIP# authors for trellis quantization and incoherence-processing foundations.
- turboderp and ExLlamaV3 contributors for EXL3/Trellis quantization and its encoder/runtime ecosystem.
- The Qwen Team for Qwen3 and
Qwen/Qwen3-30B-A3B-Base(Apache-2.0). - Z.ai for GLM-5.2, the source model in the preserved prior-control lineage.
- malaiwah for the GLM-5.2 MTP-78 overlay and calibration capture, and Josh Cartu for the associated MTP-78 recipe and rank-sliced runtime work. These credits identify the named predecessor's provenance; the Qwen experiment does not incorporate the MTP-78 runtime or draft layer.
- The MC-MoE, HIGGS, PQI, GuidedQuant, MoEQuant, EAC-MoE, and VSRAQ authors
for the mixed-precision, end-loss, router-aware, and route-shift research
identified precisely in
docs/REFERENCES.md. - Luke Alonso and Local Inference Lab contributors for B12X, the vLLM fork,
and local runtime research. The official BTX writer ports atom assembly from
B12X
btx_synth.pyat the pinned commit; that module is marked Apache-2.0 and the complete upstream license is preserved inTHIRD_PARTY_LICENSES/B12X-APACHE-2.0.txt. - NVIDIA Model Optimizer, vLLM, Hugging Face Transformers, safetensors, and huggingface_hub contributors.
- Google DeepMind for Gemma 4, used as the planned portability model.
Any future vendored adapter must retain the upstream license and notice in the same commit that introduces the code.
These acknowledgements identify intellectual and software lineage; they do not
imply endorsement of this repository by any named person or project. See
docs/REFERENCES.md for the method-to-component mapping and primary links.