Underdog-Ternary-1.1 (DEPRICATED)

DEPRICATED: This model, based on Ternary Bonsai 2 27B is no longer being served to underdog users as it is not performant per user feedback. We found the base model to be optimized for benchmarks vs usable for users. Please choose another model from the dozens that are available.

Underdog-Ternary by Conway is a 27-billion-parameter assistant for conversation, writing, coding, and summarization. This release combines post training with compact ternary weights for local inference. We applied post training to this model to run better in Underdog for human users, based on Bonsai 2.

Choose GGUF for accelerated generation with Splash, or the MLX package for native Apple Silicon workflows.

Local performance

On an Apple M5 Max, Splash with DFlash2 delivered up to 157.9 generated tokens per second in our coding test.

Task MLX (bundled runtime) GGUF (llama.cpp) GGUF (Splash + DFlash2)
Email Writing 42.0 tokens/s 45.7 tokens/s 78.2 tokens/s
Code Generation 41.0 tokens/s 45.4 tokens/s 157.9 tokens/s
Long-Text Summarization 40.4 tokens/s 44.7 tokens/s 75.4 tokens/s

Measured on Apple M5 Max with 40 GPU cores and 48 GiB unified memory. Both configurations use the same checkpoint in their respective MLX and GGUF packages. Values are means of two runs per task with temperature 0, thinking disabled, a 512-token output limit, and no reused prompt prefix. Engines were measured separately. Input lengths were 43, 44, and 3,003 tokens; generated output lengths may differ. Model loading and warmup are excluded. Splash 1.1.0 uses DFlash2 and INT8 KV; MLX 0.32.0 / MLX-LM 0.31.3 uses the bundled Hadamard-aware runtime without a draft model. The table reports generation throughput.

Run with Splash

On an Apple Silicon Mac, install Splash 1.1.0 or later, then run:

splash serve \
  --model ConwayResearch/Underdog-Ternary-1.1:PQ2_0 \
  --language-only \
  --default-reasoning-effort medium

The explicit :PQ2_0 selector selects the GGUF; Splash prepares its compatible draft model automatically. Open the chat page shown by the service, or connect your application through its OpenAI-compatible API. For medium thinking in the chat page, select Medium in the thinking menu.

Run GGUF with llama.cpp on Windows (NVIDIA)

The ternary GGUF (PQ2_0) runs on PrismML's llama.cpp build. Download its latest Windows CUDA release from the releases page (llama-prism-*-bin-win-cuda-*-x64.zip and the matching cudart-llama-bin-win-cuda-*-x64.zip, extracted into one folder). Then download the GGUF from this repository and the DFlash2 draft model:

hf download ConwayResearch/Underdog-Ternary-1.1 Underdog-PQ2_0.gguf --local-dir ./underdog-2bit
hf download incoai/Qwen3.8-27B-DFlash2-GGUF Qwen3.8-27B-DFlash2-Q8_0.gguf --local-dir ./underdog-2bit
llama-server.exe -m .\underdog-2bit\Underdog-PQ2_0.gguf `
  -md .\underdog-2bit\Qwen3.8-27B-DFlash2-Q8_0.gguf `
  --spec-type draft-dflash --spec-draft-n-max 3 `
  -ngl 99 -ngld 99 -fa on -c 16384 --jinja

Choose your model package

Variant Files to download How to use
GGUF Underdog-PQ2_0.gguf and LICENSE Load with Splash or a compatible PQ2_0 engine.
MLX The complete mlx/ directory and LICENSE Load the mlx/ directory with its included Hadamard-aware runtime.

Download only the variant used by your application. Keep the MLX configuration, tokenizer, chat template, and runtime together with its weights.

hf download ConwayResearch/Underdog-Ternary-1.1 \
  --include "mlx/*" --include "LICENSE" \
  --local-dir ./underdog-2bit

For Python integration, the supplied MLX runtime exposes vision_artifact.load_vl_model; use load_processor=False for text inference.

Downloads last month
927
MLX
Hardware compatibility
Log In to add your hardware

2-bit

GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support