Image-Text-to-Text
GGUF
llama.cpp
qwen
amd
rocm
gfx1151
strix-halo
iu4
mtp
long-context
vision
conversational
Instructions to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: llama cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: llama cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Use Docker
docker model run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- LM Studio
- Jan
- vLLM
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- Ollama
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Ollama:
ollama run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- Unsloth Desktop
- Pi
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Docker Model Runner:
docker model run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- Lemonade
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Run and chat with the model
lemonade run user.Qwen3.8-Flash-CIRU-STRIX-IU4-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download benchmarks/v2.0/controlled-coding-42tps.json from jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4: direct link, hf CLI and curl.
- Browser
- Download file 12.3 kB
-
https://huggingface.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4/resolve/04ba81fc1c0a3a35ce76a894b8d835b0636a7815/benchmarks/v2.0/controlled-coding-42tps.json
- Command line
-
hf download hf://jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4@04ba81fc1c0a3a35ce76a894b8d835b0636a7815/benchmarks/v2.0/controlled-coding-42tps.json
-
curl -L -o controlled-coding-42tps.json https://huggingface.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4/resolve/04ba81fc1c0a3a35ce76a894b8d835b0636a7815/benchmarks/v2.0/controlled-coding-42tps.json
12.3 kB
| { | |
| "workload": "Controlled greedy short-code probe; 57 prompt tokens, 520 generated tokens, not HumanEval", | |
| "configuration": "ROCm10 fixed-six, p_min0, shortlist32768, F16targetKV, Q8draftKV,16Kcontext, performance CPU governor, greedy seed123", | |
| "rows": [ | |
| { | |
| "case": "C1", | |
| "prompt": "short-code", | |
| "tg": 42.280053297306495, | |
| "pp": 115.42901115813775, | |
| "prompt_tokens": 57, | |
| "generated_tokens": 520, | |
| "ttfp_ms": 538.7954711914062, | |
| "client_total_ms": 12814.576148986816, | |
| "client_after_first_ms": 12275.78067779541, | |
| "draft_generated": 674, | |
| "draft_accepted": 400, | |
| "token_sha256": "5115118aef25826d9d904cd30f47a28fc9aeac842457b17ff906deba653fb731", | |
| "peak_ram_used_bytes": 89481662464, | |
| "delta_ram_used_bytes": 486907904, | |
| "idle_ram_used_bytes": 88994754560, | |
| "vram_available": false, | |
| "source_row_sha256": "ec1a933cdc15ba1f288d7717e8f1d2c56e48b351d7baa5054833eae6affe3476", | |
| "timings": { | |
| "cache_n": 0, | |
| "draft_n": 674, | |
| "draft_n_accepted": 400, | |
| "predicted_ms": 12275.292, | |
| "predicted_n": 520, | |
| "predicted_per_second": 42.280053297306495, | |
| "predicted_per_token_ms": 23.651815028901733, | |
| "prompt_ms": 493.81, | |
| "prompt_n": 57, | |
| "prompt_per_second": 115.42901115813775, | |
| "prompt_per_token_ms": 8.663333333333334 | |
| }, | |
| "generation_settings": { | |
| "adaptive_decay": 0.8999999761581421, | |
| "adaptive_target": -1.0, | |
| "backend_sampling": false, | |
| "chat_format": "Content-only", | |
| "dry_allowed_length": 2, | |
| "dry_base": 1.75, | |
| "dry_multiplier": 0.0, | |
| "dry_penalty_last_n": 64, | |
| "dry_sequence_breakers": [ | |
| "\n", | |
| ":", | |
| "\"", | |
| "*" | |
| ], | |
| "dynatemp_exponent": 1.0, | |
| "dynatemp_range": 0.0, | |
| "frequency_penalty": 0.0, | |
| "generation_prompt": "", | |
| "grammar": "", | |
| "grammar_lazy": false, | |
| "grammar_triggers": [], | |
| "ignore_eos": false, | |
| "logit_bias": [], | |
| "lora": [], | |
| "max_tokens": 520, | |
| "min_keep": 0, | |
| "min_p": 0.0, | |
| "mirostat": 0, | |
| "mirostat_eta": 0.10000000149011612, | |
| "mirostat_tau": 5.0, | |
| "n_discard": 0, | |
| "n_keep": 0, | |
| "n_predict": 520, | |
| "n_probs": 0, | |
| "post_sampling_probs": false, | |
| "presence_penalty": 0.0, | |
| "preserved_tokens": [], | |
| "reasoning_format": "deepseek", | |
| "reasoning_in_content": false, | |
| "repeat_last_n": 64, | |
| "repeat_penalty": 1.0, | |
| "samplers": [ | |
| "penalties", | |
| "dry", | |
| "top_n_sigma", | |
| "top_k", | |
| "typ_p", | |
| "top_p", | |
| "min_p", | |
| "xtc", | |
| "temperature" | |
| ], | |
| "seed": 123, | |
| "speculative.types": "none,draft-mtp", | |
| "stop": [], | |
| "stream": true, | |
| "temperature": 0.0, | |
| "timings_per_token": false, | |
| "top_k": 1, | |
| "top_n_sigma": -1.0, | |
| "top_p": 1.0, | |
| "typical_p": 1.0, | |
| "xtc_probability": 0.0, | |
| "xtc_threshold": 0.10000000149011612 | |
| }, | |
| "raw_sha256": "9495dee67c6af80c616aa71b5dcb9442ae452188a1d898d734c4d13969122c51", | |
| "command": { | |
| "argv": [ | |
| "/srv/llm/work/sozo-adaptive-diagnosis-20260905/candidates/fixed6-graph-boundaries/bin/llama-server", | |
| "--model", | |
| "/srv/llm/models/Qwen3.8-Flash-CIRU-STRIX-IU4/Qwen3.8-Flash-CIRU-STRIX-IU4.gguf", | |
| "--alias", | |
| "qwen38-mtp-probe", | |
| "--host", | |
| "127.0.0.1", | |
| "--port", | |
| "18186", | |
| "--jinja", | |
| "--ple-sidecar", | |
| "/srv/llm/models/Qwen3.8-Flash-CIRU-STRIX-IU4/ple", | |
| "--ple-cache-mib", | |
| "4096", | |
| "-ngl", | |
| "all", | |
| "-sm", | |
| "none", | |
| "--fit", | |
| "off", | |
| "-c", | |
| "16384", | |
| "-b", | |
| "2048", | |
| "-ub", | |
| "512", | |
| "--parallel", | |
| "1", | |
| "-t", | |
| "8", | |
| "-tb", | |
| "8", | |
| "-ctk", | |
| "f16", | |
| "-ctv", | |
| "f16", | |
| "-fa", | |
| "on", | |
| "--cont-batching", | |
| "--cache-prompt", | |
| "--cache-ram", | |
| "8192", | |
| "--cache-idle-slots", | |
| "--ctx-checkpoints", | |
| "32", | |
| "--checkpoint-min-step", | |
| "8192", | |
| "--metrics", | |
| "--slots", | |
| "--spec-type", | |
| "draft-mtp", | |
| "--spec-draft-model", | |
| "/srv/llm/models/Qwen3.8-Flash-CIRU-STRIX-IU4/mtp/Qwen3.8-Flash-CIRU-STRIX-IU4-MTP-Q8_0.gguf", | |
| "--spec-draft-ngl", | |
| "all", | |
| "--spec-draft-device", | |
| "ROCm0", | |
| "--spec-draft-type-k", | |
| "q8_0", | |
| "--spec-draft-type-v", | |
| "q8_0", | |
| "--spec-draft-threads", | |
| "8", | |
| "--spec-draft-threads-batch", | |
| "8", | |
| "--spec-draft-n-max", | |
| "6", | |
| "--spec-draft-n-min", | |
| "0", | |
| "--spec-draft-p-min", | |
| "0", | |
| "--spec-draft-p-split", | |
| "0.10" | |
| ], | |
| "env": { | |
| "GGML_CUDA_Q41_MOE_FORCE_J": "32", | |
| "GGML_QWEN4EXP_PLE_WORKERS": "16", | |
| "GGML_QWEN4EXP_PLE_STRICT_SHA": "0", | |
| "ROCBLAS_USE_HIPBLASLT": "1", | |
| "LD_LIBRARY_PATH": "/srv/llm/work/sozo-adaptive-diagnosis-20260905/candidates/fixed6-graph-boundaries/bin:/srv/llm/engines/Qwen3.8-Flash-CIRU-STRIX-IU4-v1.1/build-hip-rocm10-release/bin:/srv/llm/toolchains/therock-gfx1151-10.0.0/lib:/srv/llm/toolchains/therock-gfx1151-10.0.0/lib/rocm_sysdeps/lib:/srv/llm/toolchains/therock-gfx1151-10.0.0/lib/llvm/lib", | |
| "GGML_QSA_LONG_TOPK": "1", | |
| "GGML_QSA_RESTORE_FAST": "1", | |
| "CIRU_MTP_TOPK10": "1", | |
| "CIRU_MTP_SHORTLIST": "32768" | |
| } | |
| } | |
| }, | |
| { | |
| "case": "C2", | |
| "prompt": "short-code", | |
| "tg": 42.314176153996854, | |
| "pp": 106.23189864358639, | |
| "prompt_tokens": 57, | |
| "generated_tokens": 520, | |
| "ttfp_ms": 581.5215110778809, | |
| "client_total_ms": 12847.277402877808, | |
| "client_after_first_ms": 12265.755891799927, | |
| "draft_generated": 674, | |
| "draft_accepted": 400, | |
| "token_sha256": "5115118aef25826d9d904cd30f47a28fc9aeac842457b17ff906deba653fb731", | |
| "peak_ram_used_bytes": 89557381120, | |
| "delta_ram_used_bytes": 521732096, | |
| "idle_ram_used_bytes": 89035649024, | |
| "vram_available": false, | |
| "source_row_sha256": "afe2c7def7083fb2199d0e47da8ab0c22ad32f51de136f2596087cfc2ef91e3d", | |
| "timings": { | |
| "cache_n": 0, | |
| "draft_n": 674, | |
| "draft_n_accepted": 400, | |
| "predicted_ms": 12265.393, | |
| "predicted_n": 520, | |
| "predicted_per_second": 42.314176153996854, | |
| "predicted_per_token_ms": 23.632741811175336, | |
| "prompt_ms": 536.562, | |
| "prompt_n": 57, | |
| "prompt_per_second": 106.23189864358639, | |
| "prompt_per_token_ms": 9.413368421052631 | |
| }, | |
| "generation_settings": { | |
| "adaptive_decay": 0.8999999761581421, | |
| "adaptive_target": -1.0, | |
| "backend_sampling": false, | |
| "chat_format": "Content-only", | |
| "dry_allowed_length": 2, | |
| "dry_base": 1.75, | |
| "dry_multiplier": 0.0, | |
| "dry_penalty_last_n": 64, | |
| "dry_sequence_breakers": [ | |
| "\n", | |
| ":", | |
| "\"", | |
| "*" | |
| ], | |
| "dynatemp_exponent": 1.0, | |
| "dynatemp_range": 0.0, | |
| "frequency_penalty": 0.0, | |
| "generation_prompt": "", | |
| "grammar": "", | |
| "grammar_lazy": false, | |
| "grammar_triggers": [], | |
| "ignore_eos": false, | |
| "logit_bias": [], | |
| "lora": [], | |
| "max_tokens": 520, | |
| "min_keep": 0, | |
| "min_p": 0.0, | |
| "mirostat": 0, | |
| "mirostat_eta": 0.10000000149011612, | |
| "mirostat_tau": 5.0, | |
| "n_discard": 0, | |
| "n_keep": 0, | |
| "n_predict": 520, | |
| "n_probs": 0, | |
| "post_sampling_probs": false, | |
| "presence_penalty": 0.0, | |
| "preserved_tokens": [], | |
| "reasoning_format": "deepseek", | |
| "reasoning_in_content": false, | |
| "repeat_last_n": 64, | |
| "repeat_penalty": 1.0, | |
| "samplers": [ | |
| "penalties", | |
| "dry", | |
| "top_n_sigma", | |
| "top_k", | |
| "typ_p", | |
| "top_p", | |
| "min_p", | |
| "xtc", | |
| "temperature" | |
| ], | |
| "seed": 123, | |
| "speculative.types": "none,draft-mtp", | |
| "stop": [], | |
| "stream": true, | |
| "temperature": 0.0, | |
| "timings_per_token": false, | |
| "top_k": 1, | |
| "top_n_sigma": -1.0, | |
| "top_p": 1.0, | |
| "typical_p": 1.0, | |
| "xtc_probability": 0.0, | |
| "xtc_threshold": 0.10000000149011612 | |
| }, | |
| "raw_sha256": "e9c57e471feea6fad3b188c24dc67277d040df0772bdb6d7b86fff5b6fa6571f", | |
| "command": { | |
| "argv": [ | |
| "/srv/llm/work/sozo-adaptive-diagnosis-20260905/candidates/fixed6-graph-boundaries/bin/llama-server", | |
| "--model", | |
| "/srv/llm/models/Qwen3.8-Flash-CIRU-STRIX-IU4/Qwen3.8-Flash-CIRU-STRIX-IU4.gguf", | |
| "--alias", | |
| "qwen38-mtp-probe", | |
| "--host", | |
| "127.0.0.1", | |
| "--port", | |
| "18186", | |
| "--jinja", | |
| "--ple-sidecar", | |
| "/srv/llm/models/Qwen3.8-Flash-CIRU-STRIX-IU4/ple", | |
| "--ple-cache-mib", | |
| "4096", | |
| "-ngl", | |
| "all", | |
| "-sm", | |
| "none", | |
| "--fit", | |
| "off", | |
| "-c", | |
| "16384", | |
| "-b", | |
| "2048", | |
| "-ub", | |
| "512", | |
| "--parallel", | |
| "1", | |
| "-t", | |
| "8", | |
| "-tb", | |
| "8", | |
| "-ctk", | |
| "f16", | |
| "-ctv", | |
| "f16", | |
| "-fa", | |
| "on", | |
| "--cont-batching", | |
| "--cache-prompt", | |
| "--cache-ram", | |
| "8192", | |
| "--cache-idle-slots", | |
| "--ctx-checkpoints", | |
| "32", | |
| "--checkpoint-min-step", | |
| "8192", | |
| "--metrics", | |
| "--slots", | |
| "--spec-type", | |
| "draft-mtp", | |
| "--spec-draft-model", | |
| "/srv/llm/models/Qwen3.8-Flash-CIRU-STRIX-IU4/mtp/Qwen3.8-Flash-CIRU-STRIX-IU4-MTP-Q8_0.gguf", | |
| "--spec-draft-ngl", | |
| "all", | |
| "--spec-draft-device", | |
| "ROCm0", | |
| "--spec-draft-type-k", | |
| "q8_0", | |
| "--spec-draft-type-v", | |
| "q8_0", | |
| "--spec-draft-threads", | |
| "8", | |
| "--spec-draft-threads-batch", | |
| "8", | |
| "--spec-draft-n-max", | |
| "6", | |
| "--spec-draft-n-min", | |
| "0", | |
| "--spec-draft-p-min", | |
| "0", | |
| "--spec-draft-p-split", | |
| "0.10" | |
| ], | |
| "env": { | |
| "GGML_CUDA_Q41_MOE_FORCE_J": "32", | |
| "GGML_QWEN4EXP_PLE_WORKERS": "16", | |
| "GGML_QWEN4EXP_PLE_STRICT_SHA": "0", | |
| "ROCBLAS_USE_HIPBLASLT": "1", | |
| "LD_LIBRARY_PATH": "/srv/llm/work/sozo-adaptive-diagnosis-20260905/candidates/fixed6-graph-boundaries/bin:/srv/llm/engines/Qwen3.8-Flash-CIRU-STRIX-IU4-v1.1/build-hip-rocm10-release/bin:/srv/llm/toolchains/therock-gfx1151-10.0.0/lib:/srv/llm/toolchains/therock-gfx1151-10.0.0/lib/rocm_sysdeps/lib:/srv/llm/toolchains/therock-gfx1151-10.0.0/lib/llvm/lib", | |
| "GGML_QSA_LONG_TOPK": "1", | |
| "GGML_QSA_RESTORE_FAST": "1", | |
| "CIRU_MTP_TOPK10": "1", | |
| "CIRU_MTP_SHORTLIST": "32768" | |
| } | |
| } | |
| } | |
| ], | |
| "scope": "Two retained candidate repetitions inside order-balanced graph comparison. Raw output tokens identical. Not a native-thinking HumanEval average or a competitor speedup.", | |
| "binary_hashes_match_current_16_of_16": true, | |
| "reasoning_mode": "Explicitly closed think block before generation; non-thinking coding probe", | |
| "prompt": "<|im_start|>user\nWrite a Python TTL LRU cache using collections.OrderedDict and time.monotonic. Implement get, put, capacity eviction, and lazy expiration. Include three concise unit tests. Explain the expiry and recency invariants.<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n" | |
| } | |