HuggingFaceTB/everyday-conversations-llama3.1-2k
Viewer • Updated • 2.38k • 2.51k • 138
How to use ertghiu256/Qwen3-Hermes-4b with llama.cpp:
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M # Run inference directly in the terminal: llama cli -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M # Run inference directly in the terminal: llama cli -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
docker model run hf.co/ertghiu256/Qwen3-Hermes-4b:Q4_K_M
How to use ertghiu256/Qwen3-Hermes-4b with Ollama:
ollama run hf.co/ertghiu256/Qwen3-Hermes-4b:Q4_K_M
How to use ertghiu256/Qwen3-Hermes-4b with Pi:
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
"providers": {
"llama-cpp": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "none",
"models": [
{
"id": "ertghiu256/Qwen3-Hermes-4b:Q4_K_M"
}
]
}
}
}# Start Pi in your project directory: pi
How to use ertghiu256/Qwen3-Hermes-4b with Docker Model Runner:
docker model run hf.co/ertghiu256/Qwen3-Hermes-4b:Q4_K_M
How to use ertghiu256/Qwen3-Hermes-4b with Lemonade:
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ertghiu256/Qwen3-Hermes-4b:Q4_K_M
lemonade run user.Qwen3-Hermes-4b-Q4_K_M
lemonade list
How to use ertghiu256/Qwen3-Hermes-4b with Hermes Agent:
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ertghiu256/Qwen3-Hermes-4b:Q4_K_M
hermes
How to use ertghiu256/Qwen3-Hermes-4b with OpenClaw:
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ertghiu256/Qwen3-Hermes-4b:Q4_K_M
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ertghiu256/Qwen3-Hermes-4b:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
openclaw agent --local --agent main --message "Hello from Hugging Face"
This Qwen 3 4B model was fine-tuned on the Hermes 3 dataset to enhance its general chatting capabilities while retaining Qwen's Reasoning capabilities.
As the qwen team suggested to use
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "ertghiu256/Qwen3-Hermes-4b"
# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
# conduct text completion
generated_ids = model.generate(
**model_inputs,
max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
# parsing thinking content
try:
# rindex finding 151668 (</think>)
index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
index = 0
thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")
print("thinking content:", thinking_content)
print("content:", content)
Run this command
vllm serve ertghiu256/Qwen3-Hermes-4b --enable-reasoning --reasoning-parser deepseek_r1
Run this command
python -m sglang.launch_server --model-path ertghiu256/Qwen3-Hermes-4b --reasoning-parser deepseek-r1
Run this command
llama-server --hf-repo ertghiu256/Qwen3-Hermes-4b
or
llama-cli --hf ertghiu256/Qwen3-Hermes-4b
Run this command
ollama run hf.co/ertghiu256/Qwen3-Hermes-4b:Q4_K_M
Search
ertghiu256/Qwen3-Hermes-4b
in the lm studio model search list then download
Base model
Qwen/Qwen3-4B-Base