Instructions to use 0xKitkat/Agnes-3.0-Flash-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Use Docker
docker model run hf.co/0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "0xKitkat/Agnes-3.0-Flash-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "0xKitkat/Agnes-3.0-Flash-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
- Ollama
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with Ollama:
ollama run hf.co/0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with Docker Model Runner:
docker model run hf.co/0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
- Lemonade
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Agnes-3.0-Flash-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use 0xKitkat/Agnes-3.0-Flash-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "0xKitkat/Agnes-3.0-Flash-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download build/convert_agnes.py from 0xKitkat/Agnes-3.0-Flash-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 4.95 kB
-
https://huggingface.co/0xKitkat/Agnes-3.0-Flash-GGUF/resolve/9eae3a3f7527754ff19ceb80d862b15565e67c1e/build/convert_agnes.py
- Command line
-
hf download hf://0xKitkat/Agnes-3.0-Flash-GGUF@9eae3a3f7527754ff19ceb80d862b15565e67c1e/build/convert_agnes.py
-
curl -L -o convert_agnes.py https://huggingface.co/0xKitkat/Agnes-3.0-Flash-GGUF/resolve/9eae3a3f7527754ff19ceb80d862b15565e67c1e/build/convert_agnes.py
4.95 kB
| """Convert Agnes with an exact SwiGLU branch merge into the Qwen3.5 GGUF graph. | |
| Run with the same arguments as llama.cpp/convert_hf_to_gguf.py. | |
| The HF checkpoint and its native architecture remain unchanged. | |
| """ | |
| import copy | |
| import json | |
| import runpy | |
| import sys | |
| from pathlib import Path | |
| import torch | |
| ROOT = Path(__file__).resolve().parent | |
| LLAMA = ROOT / "llama.cpp" | |
| sys.path.insert(0, str(LLAMA)) | |
| sys.path.insert(0, str(LLAMA / "gguf-py")) | |
| import conversion | |
| from conversion.base import ModelBase, gguf | |
| from conversion.qwen import Qwen3_5TextModel | |
| from conversion.qwen3vl import Qwen3VLVisionModel | |
| class AgnesTextGGUF(Qwen3_5TextModel): | |
| model_arch = gguf.MODEL_ARCH.QWEN35 | |
| def get_vocab_base(self): | |
| from transformers import PreTrainedTokenizerFast | |
| tokenizer = PreTrainedTokenizerFast.from_pretrained(self.dir_model, fix_mistral_regex=False) | |
| pre = self.get_vocab_base_pre(tokenizer) | |
| vocab = tokenizer.get_vocab() | |
| size = self.hparams["vocab_size"] | |
| if max(vocab.values()) >= size: | |
| raise ValueError("Tokenizer exceeds embedding vocabulary") | |
| reverse = {idx: token for token, idx in vocab.items()} | |
| added = tokenizer.added_tokens_decoder | |
| tokens, types = [], [] | |
| for idx in range(size): | |
| token = reverse.get(idx, f"[PAD{idx}]") | |
| kind = gguf.TokenType.NORMAL | |
| if idx not in reverse: | |
| kind = gguf.TokenType.UNUSED | |
| elif idx in added: | |
| if not added[idx].normalized: | |
| token = tokenizer.decode(tokenizer.encode(token, add_special_tokens=False)) | |
| kind = gguf.TokenType.CONTROL if added[idx].special or self.does_token_look_special(token) else gguf.TokenType.USER_DEFINED | |
| tokens.append(token) | |
| types.append(kind) | |
| return tokens, types, pre | |
| def __init__(self, dir_model, *args, **kwargs): | |
| config = copy.deepcopy(kwargs.get("hparams") or json.loads((dir_model / "config.json").read_text())) | |
| text = config.get("text_config", config) | |
| text["intermediate_size"] += text["parallel_ffn_intermediate_size"] | |
| text["layer_types"] = [ | |
| {"agnes_delta_attention": "linear_attention", "agnes_global_attention": "full_attention"}[k] | |
| for k in text["layer_types"] | |
| ] | |
| text["full_attention_interval"] = text.get("global_attention_interval", 4) | |
| text["mtp_num_hidden_layers"] = 0 | |
| kwargs["hparams"] = config | |
| type(self).no_mtp = True | |
| super().__init__(dir_model, *args, **kwargs) | |
| def filter_tensors(cls, item): | |
| name, gen = item | |
| if name.startswith(("mtp.", "model.mtp.")): | |
| return None | |
| name = name.replace(".delta_attn.", ".linear_attn.").replace(".global_attn.", ".self_attn.") | |
| return super().filter_tensors((name, gen)) | |
| def index_tensors(self, remote_hf_model_id=None): | |
| tensors = super().index_tensors(remote_hf_model_id) | |
| branches = [name for name in tensors if ".mlp.parallel_ffn." in name] | |
| if len(branches) != 3 * self.hparams["text_config"]["num_hidden_layers"]: | |
| raise ValueError("Expected exactly three parallel MLP matrices per decoder layer") | |
| for branch in branches: | |
| main = branch.replace(".parallel_ffn", "") | |
| if main not in tensors: | |
| raise ValueError(f"Missing main branch: {main}") | |
| a, b = tensors[main], tensors.pop(branch) | |
| dim = 1 if main.endswith("down_proj.weight") else 0 | |
| tensors[main] = lambda a=a, b=b, dim=dim: torch.cat((a(), b()), dim=dim) | |
| return tensors | |
| def get_vocab_base_pre(self, tokenizer): | |
| backend = json.loads(tokenizer.backend_tokenizer.to_str()) | |
| expected = r"(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\r\n\p{L}\p{N}]?[\p{L}\p{M}]+|\p{N}| ?[^\s\p{L}\p{M}\p{N}]+[\r\n]*|\s*[\r\n]+|\s+(?!\S)|\s+" | |
| sequence = backend["pre_tokenizer"] | |
| if sequence["type"] != "Sequence": | |
| raise ValueError("Unexpected pre-tokenizer structure") | |
| steps = sequence["pretokenizers"] | |
| if len(steps) != 2 or steps[0].get("pattern", {}).get("Regex") != expected: | |
| raise ValueError("Agnes pre-tokenizer differs from verified Qwen3.5 regex") | |
| if steps[1]["type"] != "ByteLevel" or steps[1].get("add_prefix_space") or steps[1].get("use_regex"): | |
| raise ValueError("Unexpected ByteLevel settings") | |
| return "qwen35" | |
| class AgnesVisionGGUF(Qwen3VLVisionModel): | |
| model_arch = Qwen3VLVisionModel.model_arch | |
| conversion.TEXT_MODEL_MAP["AgnesForConditionalGeneration"] = "qwen" | |
| conversion.MMPROJ_MODEL_MAP["AgnesForConditionalGeneration"] = "qwen3vl" | |
| if __name__ == "__main__": | |
| runpy.run_path(str(LLAMA / "convert_hf_to_gguf.py"), run_name="__main__") | |