Instructions to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Use Docker
docker model run hf.co/lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
- Ollama
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with Ollama:
ollama run hf.co/lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with Docker Model Runner:
docker model run hf.co/lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
- Lemonade
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Apriel-Nemotron-15b-Thinker-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "lmstudio-community/Apriel-Nemotron-15b-Thinker-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Wrong jinja Template
I get this error
Failed to parse Jinja template: Parser Error: Expected closing statement token. Comma !== CloseStatement.
Would be nice if you could fix that (;
This Template works:
{{- bos_token }}
{%- if messages[0]['role'] == 'system' %}
{%- set system_message = messages[0]['content']|trim %}
{%- set messages = messages[1:] %}
{%- else %}
{%- set system_message = "" %}
{%- endif %}
{{- "<|start_header_id|>system<|end_header_id|>\n\n" }}{{- system_message }}{{- "<|eot_id|>" }}
{%- for message in messages %}
{%- if message['role'] == 'assistant' and '' in message['content'] %}
{%- set content = message['content'].split('')|last %}
{%- else %}
{%- set content = message['content'] %}
{%- endif %}
{{- '<|start_header_id|>' + message['role'] + '<|end_header_id|>\n\n' + content | trim + '<|eot_id|>' }}
{%- endfor %}
{%- if add_generation_prompt %}
{{- '<|start_header_id|>assistant<|end_header_id|>\n\n' }}
{%- endif %}
Thank you!
There are couple of issues with this template replacement:
- Tool calling is gone
- Pre-defined system prompt is gone (this is a major issue, because it is highly recommended to use certain predefined system prompt with this model and add all user instructions into the prompt itself)
- Tags are similar, but not identical to the ones used in the original template, possibly leading to issues with the model
All of these issues could lead to sub-optimal performance of the model.
I manually identified the part of the template code which was incompatible with LM Studio. The issue was in the tool calling part, starting with set. I turned to Grok for help, pointed at the part of the code where the issue was and the AI helped me to fix the template as follows:
{%- set add_tool_id = true -%}
{%- set reasoning_prompt = "You are a thoughtful and systematic AI assistant built by ServiceNow Language Models (SLAM) lab. Before providing an answer, analyze the problem carefully and present your reasoning step by step. After explaining your thought process, provide the final solution in the following format: [BEGIN FINAL RESPONSE] ... [END FINAL RESPONSE]." -%}
{%- set reasoning_asst_turn_start = "Here are my reasoning steps:\n" -%}
{%- if messages[0]["role"] != "system" -%}
{%- if tools is defined and tools is not none and tools | length > 0 -%}
{{- "<|system|>\n" + reasoning_prompt + "\n\nYou are provided with function signatures within <available_tools></available_tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about the arguments. You should infer the argument values from previous user responses and the system message. Here are the available tools: <available_tools>" -}}
{%- for tool in tools -%}
{{- tool -}}
{%- endfor -%}
{{- "</available_tools>.\n\nReturn all function calls as a list of json objects within <tool_calls></tool_calls> XML tags. Each json object should contain a function name and arguments as follows: <tool_calls>[{\"name\": <function-name-1>, \"arguments\": <args-dict-1>}, {\"name\": <function-name-2>, \"arguments\": <args-dict-2>},...]</tool_calls>\n<|end|>\n" -}}
{%- else -%}
{{- "<|system|>\n" + reasoning_prompt + "\n<|end|>\n" -}}
{%- endif -%}
{%- endif -%}
{%- for message in messages -%}
{%- if message["role"] == "user" -%}
{{- "<|user|>\n" + message["content"] + "\n<|end|>\n" -}}
{%- elif message["role"] == "system" -%}
{%- if tools is defined and tools is not none and tools | length > 0 -%}
{{- "<|system|>\n" + reasoning_prompt + "\n\n" + message["content"] + "\nYou are provided with function signatures within <available_tools></available_tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about the arguments. You should infer the argument values from previous user responses and the system message. Here are the available tools: <available_tools>" -}}
{%- for tool in tools -%}
{{- tool -}}
{%- endfor -%}
{{- "</available_tools>.\n\nReturn all function calls as a list of json objects within <tool_calls></tool_calls> XML tags. Each json object should contain a function name and arguments as follows: <tool_calls>[{\"name\": <function-name-1>, \"arguments\": <args-dict-1>}, {\"name\": <function-name-2>, \"arguments\": <args-dict-2>},...]</tool_calls>\n<|end|>\n" -}}
{%- else -%}
{{- "<|system|>\n" + reasoning_prompt + "\n\n" + message["content"] + "\n<|end|>\n" -}}
{%- endif -%}
{%- elif message["role"] == "assistant" -%}
{%- if loop.last -%}
{%- set add_tool_id = false -%}
{%- endif -%}
{{- "<|assistant|>\n" -}}
{%- if message["content"] is not none -%}
{{- message["content"] -}}
{%- elif message["chosen"] is not none and message["chosen"] | length > 0 -%}
{{- message["chosen"][0] -}}
{%- endif -%}
{%- if message["tool_calls"] is not none and message["tool_calls"] | length > 0 -%}
{{- "\n<tool_calls>[" -}}
{%- for tool_call in message["tool_calls"] -%}
{{- "{\"name\": \"" + tool_call["function"]["name"] + "\", \"arguments\": " + tool_call["function"]["arguments"] | string -}}
{%- if add_tool_id -%}
{{- ", \"id\": \"" + tool_call["id"] -}}
{%- endif -%}
{{- "}" -}}
{%- if not loop.last -%}
{{- ", " -}}
{%- endif -%}
{%- endfor -%}
{{- "]</tool_calls>" -}}
{%- endif -%}
{%- if not loop.last or training_prompt -%}
{{- "\n<|end|>\n" -}}
{%- endif -%}
{%- elif message["role"] == "tool" -%}
{{- "<|tool_result|>\n" + message["content"] | string + "\n<|end|>\n" -}}
{%- endif -%}
{%- if loop.last and add_generation_prompt and message["role"] != "assistant" -%}
{{- "<|assistant|>\n" + reasoning_asst_turn_start -}}
{%- endif -%}
{%- endfor -%}
This template should retain the original functionality and it works in LM Studio (I just tested it myself).