Text Generation
GGUF
llama.cpp
qwen3.8
conversational
roleplay
creative-writing
character
humanlike
uncensored
sillytavern
imatrix
Instructions to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Use Docker
docker model run hf.co/0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
- Ollama
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with Ollama:
ollama run hf.co/0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with Docker Model Runner:
docker model run hf.co/0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
- Lemonade
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3.8-27B-Humanlike-Chat-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "0xSojalSec/Qwen3.8-27B-Humanlike-Chat-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 4,769 Bytes
f1f655f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 | # Step-863 Artifact and Reasoning Findings
Historical step-863 characterization as of 2026-09-04, not results for the current step-576 strength-0.7 rebuild. Current artifact lineage, public WikiText calibration and release gates are recorded in [`artifact-manifest.json`](artifact-manifest.json). No private conversation text is included here.
## Artifact lineage
The corrected GGUF path uses the pinned Huihui Qwen3.8-27B base and step-863 rank-256, alpha-32 LoRA. The adapter conversion includes the required value-head reorder correction. The original merged F16 matched the approved hot-LoRA eight-turn non-thinking branch exactly in that earlier sampler profile.
All quant tiers derive directly from the same verified F16 master, never by requantizing a lower-bit file. Q6/Q5/Q4/Q3 use the shared 61,173-token activation-calibration corpus's importance matrix, covering 496 tensors, with 96 recurrent gate weights protected at Q8_0. Calibration requested balanced off/low/medium render strata, but off and medium complete-chat renderings were identical; it did not simulate generated reasoning trajectories.
The new BF16 companion is separately merged from BF16 base plus FP32 adapter with a direct F32-to-BF16 output cast. Non-thinking qualification is separate from the older F16 reasoning characterization below.
## Reasoning characterization
Three recorded seeds (42, 31415, 271828), low and medium reasoning, eight-turn generated-history branches, no system prompt, no inference retries. Each branch used its own generated finals; strict reasoning-only EOS outputs were retained as no-answer and omitted from assistant history. Turns within a branch are correlated.
| Artifact | Recorded turns | Valid finals | Strict reasoning-only no-answers | Other invalid rows |
|---|---:|---:|---:|---:|
| Merged F16 | 48 | 18 | 30 | 0 |
| Q8_0 | 45 of 48 planned | 12 | 32 | 1 |
| Q6_K | 48 | 15 | 33 | 0 |
| Q5_K_M | 48 | 13 | 35 | 0 |
Q8 low/seed-271828 stopped at turn five because sampled EOS was outside the captured top-10 alternatives, triggering the frozen token-evidence guard. That five-row capture remains partial; the three unrun turns are not treated as observations or successes. The other independent branches were completed without retrying this branch.
The raw failing generations ended before emitting the reasoning-close token and a final answer. They were not merely stripped by a downstream parser, and they stopped on EOS rather than exhausting the token cap.
## Matched first-turn diagnosis
| Arm | Valid thinking finals |
|---|---:|
| Unadapted BF16 base, stock llama.cpp | 6/6 |
| BF16 base plus FP32 LoRA, stock llama.cpp | 3/6 |
| Merged F16, stock llama.cpp | 3/6 |
| Merged F16 plus explicit reasoning-boundary instruction | 3/6 |
| Merged F16 plus reference GDN normalization patch | 3/6 |
Repeating the six F16 first-turn controls with the corrected neutral thinking-penalty sampler reproduced the original prompt hashes, raw outputs and token sequences exactly. The tested prompt instruction and GDN patch did not rescue the failure. The patch is not the release runtime and is not claimed as a fix.
## Serverless comparison
A bounded live vLLM/serverless check confirmed identical low-mode prompt token IDs and two successful reasoning-plus-final responses at seeds 42 and 31415. The latter seed fails in the matched llama.cpp controls. The third request was canceled at the diagnostic deadline; medium was not reached. Do not interpret this as a completed six-cell serverless result.
## Sampling correction
The earlier evaluator reported non-thinking presence penalty 1.5 but omitted the `penalties` sampler, so that nonzero penalty was inactive. The source was corrected and the original evaluator preserved. The release's non-thinking checks use the active penalties sampler. Thinking penalties were already neutral, and all 141 newly recorded thinking requests matched the intended numeric settings, so this error does not explain their missing finals.
## Operating recommendation
Use explicit thinking-off settings. Release verification is limited to artifact integrity, the private first-message gate and a short non-thinking replay. Do not infer broad conversational quality, reasoning reliability, or cross-runtime equivalence from these checks.
## Primary sources
- [Qwen3.8-27B sampling recommendations](https://huggingface.co/Qwen/Qwen3.8-27B#best-practices)
- [Pinned llama.cpp sampler implementation](https://github.com/ggml-org/llama.cpp/blob/95ef7fc16054e63b427a3ef00188e055ef7586d8/common/sampling.cpp)
- [Proposed GDN normalization correction](https://github.com/ggml-org/llama.cpp/pull/28068)
- [vLLM experimental GGUF plugin and tested coverage](https://github.com/vllm-project/vllm-gguf-plugin)
|