Instructions to use ConwayResearch/Underdog-Ternary-1.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use ConwayResearch/Underdog-Ternary-1.1 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("ConwayResearch/Underdog-Ternary-1.1") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ConwayResearch/Underdog-Ternary-1.1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0 # Run inference directly in the terminal: llama cli -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0 # Run inference directly in the terminal: llama cli -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0 # Run inference directly in the terminal: ./llama-cli -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ConwayResearch/Underdog-Ternary-1.1:Q2_0
Use Docker
docker model run hf.co/ConwayResearch/Underdog-Ternary-1.1:Q2_0
- LM Studio
- Jan
- vLLM
How to use ConwayResearch/Underdog-Ternary-1.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ConwayResearch/Underdog-Ternary-1.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ConwayResearch/Underdog-Ternary-1.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ConwayResearch/Underdog-Ternary-1.1:Q2_0
- Ollama
How to use ConwayResearch/Underdog-Ternary-1.1 with Ollama:
ollama run hf.co/ConwayResearch/Underdog-Ternary-1.1:Q2_0
- Unsloth Desktop
- Pi
How to use ConwayResearch/Underdog-Ternary-1.1 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ConwayResearch/Underdog-Ternary-1.1"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ConwayResearch/Underdog-Ternary-1.1" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use ConwayResearch/Underdog-Ternary-1.1 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "ConwayResearch/Underdog-Ternary-1.1"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "ConwayResearch/Underdog-Ternary-1.1" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ConwayResearch/Underdog-Ternary-1.1", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use ConwayResearch/Underdog-Ternary-1.1 with Docker Model Runner:
docker model run hf.co/ConwayResearch/Underdog-Ternary-1.1:Q2_0
- Lemonade
How to use ConwayResearch/Underdog-Ternary-1.1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ConwayResearch/Underdog-Ternary-1.1:Q2_0
Run and chat with the model
lemonade run user.Underdog-Ternary-1.1-Q2_0
List all available models
lemonade list
- Hermes Agent
How to use ConwayResearch/Underdog-Ternary-1.1 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ConwayResearch/Underdog-Ternary-1.1"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ConwayResearch/Underdog-Ternary-1.1
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ConwayResearch/Underdog-Ternary-1.1 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "ConwayResearch/Underdog-Ternary-1.1"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ConwayResearch/Underdog-Ternary-1.1" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Underdog-Ternary-1.1 (DEPRICATED)
DEPRICATED: This model, based on Ternary Bonsai 2 27B is no longer being served to underdog users as it is not performant per user feedback. We found the base model to be optimized for benchmarks vs usable for users. Please choose another model from the dozens that are available.
Underdog-Ternary by Conway is a 27-billion-parameter assistant for conversation, writing, coding, and summarization. This release combines post training with compact ternary weights for local inference. We applied post training to this model to run better in Underdog for human users, based on Bonsai 2.
Choose GGUF for accelerated generation with Splash, or the MLX package for native Apple Silicon workflows.
Local performance
On an Apple M5 Max, Splash with DFlash2 delivered up to 157.9 generated tokens per second in our coding test.
| Task | MLX (bundled runtime) | GGUF (llama.cpp) | GGUF (Splash + DFlash2) |
|---|---|---|---|
| Email Writing | 42.0 tokens/s | 45.7 tokens/s | 78.2 tokens/s |
| Code Generation | 41.0 tokens/s | 45.4 tokens/s | 157.9 tokens/s |
| Long-Text Summarization | 40.4 tokens/s | 44.7 tokens/s | 75.4 tokens/s |
Measured on Apple M5 Max with 40 GPU cores and 48 GiB unified memory. Both configurations use the same checkpoint in their respective MLX and GGUF packages. Values are means of two runs per task with temperature 0, thinking disabled, a 512-token output limit, and no reused prompt prefix. Engines were measured separately. Input lengths were 43, 44, and 3,003 tokens; generated output lengths may differ. Model loading and warmup are excluded. Splash 1.1.0 uses DFlash2 and INT8 KV; MLX 0.32.0 / MLX-LM 0.31.3 uses the bundled Hadamard-aware runtime without a draft model. The table reports generation throughput.
Run with Splash
On an Apple Silicon Mac, install Splash 1.1.0 or later, then run:
splash serve \
--model ConwayResearch/Underdog-Ternary-1.1:PQ2_0 \
--language-only \
--default-reasoning-effort medium
The explicit :PQ2_0 selector selects the GGUF; Splash prepares its compatible draft model automatically. Open the chat page shown by the service, or connect your application through its OpenAI-compatible API. For medium thinking in the chat page, select Medium in the thinking menu.
Run GGUF with llama.cpp on Windows (NVIDIA)
The ternary GGUF (PQ2_0) runs on PrismML's llama.cpp build. Download its latest Windows CUDA release from the releases page (llama-prism-*-bin-win-cuda-*-x64.zip and the matching cudart-llama-bin-win-cuda-*-x64.zip, extracted into one folder). Then download the GGUF from this repository and the DFlash2 draft model:
hf download ConwayResearch/Underdog-Ternary-1.1 Underdog-PQ2_0.gguf --local-dir ./underdog-2bit
hf download incoai/Qwen3.8-27B-DFlash2-GGUF Qwen3.8-27B-DFlash2-Q8_0.gguf --local-dir ./underdog-2bit
llama-server.exe -m .\underdog-2bit\Underdog-PQ2_0.gguf `
-md .\underdog-2bit\Qwen3.8-27B-DFlash2-Q8_0.gguf `
--spec-type draft-dflash --spec-draft-n-max 3 `
-ngl 99 -ngld 99 -fa on -c 16384 --jinja
Choose your model package
| Variant | Files to download | How to use |
|---|---|---|
| GGUF | Underdog-PQ2_0.gguf and LICENSE |
Load with Splash or a compatible PQ2_0 engine. |
| MLX | The complete mlx/ directory and LICENSE |
Load the mlx/ directory with its included Hadamard-aware runtime. |
Download only the variant used by your application. Keep the MLX configuration, tokenizer, chat template, and runtime together with its weights.
hf download ConwayResearch/Underdog-Ternary-1.1 \
--include "mlx/*" --include "LICENSE" \
--local-dir ./underdog-2bit
For Python integration, the supplied MLX runtime exposes vision_artifact.load_vl_model; use load_processor=False for text inference.
- Downloads last month
- 927