Text Generation
Transformers
Safetensors
English
qwen3
guardpoint
valiant
valiant-labs
qwen
qwen-3
qwen-3-14b
14b
reasoning
science
science-reasoning
medicine
internal-medicine
clinical-diagnosis
medical-understanding
medical-reasoning
medical-diagnosis
medical-management
problem-solving
anatomy
angiology
bariatric
cardiovascular
dental
dermatology
endocrinology
ENT
hematology
immunology
infectious-disease
musculoskeletal
neurology
obstetrics
ophtamology
oncology
orthopedics
pathology
psychiatry
pulmonology
radiology
surgery
triage
urology
analytical
data
data-interpretation
expert
rationality
conversational
chat
instruct
text-generation-inference
Instructions to use ValiantLabs/Qwen3-14B-Guardpoint with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ValiantLabs/Qwen3-14B-Guardpoint with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ValiantLabs/Qwen3-14B-Guardpoint") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ValiantLabs/Qwen3-14B-Guardpoint") model = AutoModelForCausalLM.from_pretrained("ValiantLabs/Qwen3-14B-Guardpoint", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ValiantLabs/Qwen3-14B-Guardpoint with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ValiantLabs/Qwen3-14B-Guardpoint" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ValiantLabs/Qwen3-14B-Guardpoint", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ValiantLabs/Qwen3-14B-Guardpoint
- SGLang
How to use ValiantLabs/Qwen3-14B-Guardpoint with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ValiantLabs/Qwen3-14B-Guardpoint" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ValiantLabs/Qwen3-14B-Guardpoint", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ValiantLabs/Qwen3-14B-Guardpoint" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ValiantLabs/Qwen3-14B-Guardpoint", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ValiantLabs/Qwen3-14B-Guardpoint with Docker Model Runner:
docker model run hf.co/ValiantLabs/Qwen3-14B-Guardpoint
File size: 5,710 Bytes
7237d47 f34d11e 6f4f360 7237d47 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 | ---
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- guardpoint
- valiant
- valiant-labs
- qwen
- qwen-3
- qwen-3-14b
- 14b
- reasoning
- science
- science-reasoning
- medicine
- internal-medicine
- clinical-diagnosis
- medical-understanding
- medical-reasoning
- medical-diagnosis
- medical-management
- problem-solving
- anatomy
- angiology
- bariatric
- cardiovascular
- dental
- dermatology
- endocrinology
- ENT
- hematology
- immunology
- infectious-disease
- musculoskeletal
- neurology
- obstetrics
- ophtamology
- oncology
- orthopedics
- pathology
- psychiatry
- pulmonology
- radiology
- surgery
- triage
- urology
- analytical
- data
- data-interpretation
- expert
- rationality
- conversational
- chat
- instruct
base_model: Qwen/Qwen3-14B
datasets:
- sequelbox/Superpotion-DeepSeek-V3.2-Speciale
license: apache-2.0
---
**[Support our open-source dataset and model releases!](https://huggingface.co/spaces/sequelbox/SupportOpenSource)**

Guardpoint: [gemma-4-12B](https://huggingface.co/ValiantLabs/gemma-4-12B-it-Guardpoint), [Qwen3-14B](https://huggingface.co/ValiantLabs/Qwen3-14B-Guardpoint), [gpt-oss-20b](https://huggingface.co/ValiantLabs/gpt-oss-20b-Guardpoint), [Qwen3.5-27B](https://huggingface.co/ValiantLabs/Qwen3.5-27B-Guardpoint), [Qwen3-32B](https://huggingface.co/ValiantLabs/Qwen3-32B-Guardpoint), [gpt-oss-120b](https://huggingface.co/ValiantLabs/gpt-oss-120b-Guardpoint)
Guardpoint is a medical reasoning specialist built on Qwen 3.
- Finetuned on our high-difficulty [medical reasoning](https://huggingface.co/datasets/sequelbox/Superpotion-DeepSeek-V3.2-Speciale) data generated with [Deepseek V3.2 Speciale!](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale)
- Structured medical reasoning: organized, informative responses for medical diagnosis, management, knowledge, and understanding!
- Cut token costs: organized, concise responses use less tokens for faster inference!
- Trained on a wide variety of medical disciplines, patient profiles, and question types!
## Prompting Guide
Guardpoint delivers structured medical responses using the [Qwen 3](https://huggingface.co/Qwen/Qwen3-14B) prompt format.
Guardpoint is a reasoning finetune; **we recommend enable_thinking=True for all chats.**
Example inference script to get started:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "ValiantLabs/Qwen3-14B-Guardpoint"
# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
# prepare the model input
prompt = "A 60-year-old undergoes a Total Knee Arthroplasty (TKA). Post-operatively, they complain of a clunking sensation and instability when descending stairs. On exam, they have excessive posterior translation of the tibia at 90 degrees of flexion. The TKA used a Cruciate Retaining (CR) implant. Diagnosis is PCL incompetence or rupture. Explain why a CR implant relies on a functional PCL for femoral rollback and how converting to a Posterior Stabilized (PS) implant resolves this biomechanical failure."
#prompt = "I have that tube in my chest for dialysis while my arm heals. The dressing came off in the shower and the tube got tugged a bit. It didn't come out, but now there's this red cuff thing showing that used to be inside the skin. It’s sticking out about an inch. Can I just push it back in and tape it?"
#prompt = "In the workup of a tumor of unknown primary, a biopsy shows a poorly differentiated carcinoma. The IHC profile is: CK7+, CK20+, CDX2+, TTF-1 negative, PAX8 negative. Based on this cytokeratin and transcription factor profile, where is the most likely primary site of the malignancy?"
#prompt = "I have bad arthritis in my lower back and hips. I saw a chiropractor who said my 'pelvis is twisted' and wants to do high-velocity adjustments. My rheumatologist said absolutely not because of my 'osteophytes'. Who is right? I just want to walk without stiffness."
messages = [
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
# conduct text completion
generated_ids = model.generate(
**model_inputs,
max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()
# parsing thinking content
try:
# rindex finding 151668 (</think>)
index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
index = 0
thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")
print("thinking content:", thinking_content)
print("content:", content)
```
DISCLAIMER: Guardpoint is a medical reasoning finetune that is subject to the strengths and weaknesses of LLMs. A conversation with an LLM is not a substitute for a professional medical examination. Utilize Guardpoint responsibly.

Guardpoint is created by [Valiant Labs.](http://valiantlabs.ca/)
[Check out our HuggingFace page to see Shining Valiant, Esper, and all of our models!](https://huggingface.co/ValiantLabs)
We care about open source. For everyone to use.
|