Ling 3.0 flash chat template
A drop-in Jinja chat template for Ling 3.0 flash,
rendering the model's native conversation format. It carries the same agentic hardening as the Qwen fork
(https://huggingface.co/sanjxz/Qwen-Fixed-Chat-Templates/tree/main), re-expressed in Ling's <system> / <user> / <assistant> token scheme.
- Engines: llama.cpp / LM Studio / MLX / vLLM and any engine with a HuggingFace-Jinja (minja) runtime.
Files
| File | Internal version | What it is |
|---|---|---|
chat_template_ling.jinja |
ling-fixed |
The maintained template โ full loop-guard, security hardening, tool + skill integration, developer-role support. Use this. |
chat_template_ling-orig.jinja |
ling v3 iteration |
The minimal upstream original. Kept for reference/diffing only. |
Feature set
The agentic internals are (almost, WIP) identical to the Qwen fork โ see the "What the fork adds over the
original" section of README.md for the full write-up. In brief, this template adds
over chat_template_laguna-orig.jinja:
- Escalating loop-guard โ consecutive-failure escalation (nudge โ warn โ hard halt), plus
tool suspension: on hard halt a pre-scan replaces the
<available_tools>block with a suspension notice so the model cannot emit a tool call. - Repeat / ping-pong detection โ fingerprints each turn by tool name + arguments over a two-turn window (catches A-B-A-B alternation); mutating tools exempted from result-repeat, caught instead by the ping-pong nudge.
- Four-tier failure heuristics โ hard / harness / generic / permission errors, kept in sync
between the pre-scan and the main loop, with
$(shell) andtook(timing) false-positive escapes. Read tools are exempt from error-text matching; results are attributed viatool_call_idโ name. - Security hardening โ only system/developer messages may toggle thinking; user content has
the tokens stripped but never applied.
sanitize_tool_tokensstrips literal<|think_off|>/<|think_on|>from untrusted tool output. - Harness integration โ structural
skill-tool triggering, vanished-tool warnings,strftime_nowdate injection, OpenAI{"type":"function",โฆ}wrapper unwrapping, middle-out tool-response truncation,<__media__>marker stripping, unclosed-reasoning recovery.
Tool call formats
History tool calls and the format reminder are rendered in one of two formats, selected with tool_call_format.
XML (default) โ arguments as <arg_key>/<arg_value> pairs:
<tool_call>get_weather<arg_key>location</arg_key><arg_value>NYC</arg_value></tool_call>
Running with llama.cpp
Point --chat-template-file at the template and enable --jinja. Example for Laguna S 2.1 APEX:
import os
import subprocess
import sys
env = os.environ.copy()
cmd = [
r".\llama-server.exe",
"-m", r"D:\Ling-3.0-flash-AD-IQ3_S-00001-of-00002.gguf",
"--reasoning-preserve",
"--reasoning-budget", "16000",
"--reasoning-budget-message", "I've thought enough. Answering now with what I have.",
"--fit", "on",
"--n-cpu-moe", "39",
"-lm", "mlock", #recent llama.cpp only ; --mlock & --no-mmap alternative
"-c", "120000",
"--cache-type-k", "q8_0",
"--cache-type-v", "q8_0",
"-np", "1",
"-fa", "on",
"-t", "8",
"-tb", "8",
"-b", "2048",
"-ub", "2048",
"--jinja",
"-kvu",
"--temp", "0.6",
"--top-p", "0.95",
"--top-k", "20",
"--samplers", "top_k;top_p;temperature",
"--alias", "laguna-s-2.1",
"--cache-reuse", "256",
"--cache-ram", "1024",
"--host", "127.0.0.1",
"--port", "8080",
"--verbosity", "4",
"--chat-template-file", r"G:\xlam3\chat_template_ling.jinja",
]
if __name__ == "__main__":
try:
print("Starting llama-server...")
subprocess.run(cmd, env=env, check=True)
except KeyboardInterrupt:
print("\nServer stopped by user.")
except Exception as e:
print(f"\nError running server: {e}")
Authorship & license
| Role | Author |
|---|---|
| Base model | Poolside (Laguna S 2.1) |
| Template lineage / fixes | froggeric + fork contributors |
MIT.
Model tree for sanjxz/Ling-3.0-Flash-Agentic-Chat-Template-Jinja
Base model
froggeric/Qwen-Fixed-Chat-Templates