Instructions to use distil-labs/distil-siemens-s7-1200-docs-llama-1b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use distil-labs/distil-siemens-s7-1200-docs-llama-1b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="distil-labs/distil-siemens-s7-1200-docs-llama-1b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("distil-labs/distil-siemens-s7-1200-docs-llama-1b") model = AutoModelForCausalLM.from_pretrained("distil-labs/distil-siemens-s7-1200-docs-llama-1b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use distil-labs/distil-siemens-s7-1200-docs-llama-1b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "distil-labs/distil-siemens-s7-1200-docs-llama-1b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "distil-labs/distil-siemens-s7-1200-docs-llama-1b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/distil-labs/distil-siemens-s7-1200-docs-llama-1b
- SGLang
How to use distil-labs/distil-siemens-s7-1200-docs-llama-1b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "distil-labs/distil-siemens-s7-1200-docs-llama-1b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "distil-labs/distil-siemens-s7-1200-docs-llama-1b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "distil-labs/distil-siemens-s7-1200-docs-llama-1b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "distil-labs/distil-siemens-s7-1200-docs-llama-1b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use distil-labs/distil-siemens-s7-1200-docs-llama-1b with Docker Model Runner:
docker model run hf.co/distil-labs/distil-siemens-s7-1200-docs-llama-1b
Distil Siemens S7-1200 Docs — Llama 3.2 1B
A fine-tuned Llama 3.2 1B Instruct model, distilled for question-answering over Siemens SIMATIC S7-1200 PLC documentation. Designed to be paired with a RAG pipeline so that dense technical manuals — alarm codes, signal addresses, parameter tables — are accessible at the point of use, even on edge hardware without GPU.
Why this model
Industrial environments with strict security requirements (e.g. Purdue model network segmentation) cannot easily call cloud LLMs, and hosting large open-source models on-prem requires expensive GPUs. This model demonstrates that fine-tuned small language models offer a practical alternative: a 1B parameter model that, after distillation, matches or exceeds a 3B base model on domain-specific technical QA.
Evaluation
All models were evaluated on 144 held-out questions from the S7-1200 system manual using an LLM-as-a-Judge binary score.
| Model | Parameters | LLM-as-a-Judge Pass Rate |
|---|---|---|
| Llama 3.2 1B Instruct (base) | 1B | 45.1% |
| Llama 3.2 3B Instruct (base) | 3B | 60.4% |
| Llama 3.2 1B Instruct (this model) | 1B | 61.1% |
Fine-tuning improved the 1B model by +16 percentage points, bringing it to parity with the 3x larger 3B base model. On 6 out of 144 test questions, this model answered correctly where both the 1B and 3B base models failed.
Usage
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "distillabs/distil-siemens-s7-1200-docs-llama-1b"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
system_prompt = """You are a problem solving model working on task_description XML block:
<task_description>Answer technical questions about an industrial automation system (The S7-1200 Programmable controller) using information provided in the context passage. Make sure to provide answers that are complete and include all relevant details from the context; do not miss critical information from the context.</task_description>
You will be given a single question and a context passage. Answer the question based on the context."""
# In a RAG pipeline, `context` comes from your retriever
context = "The maximum cold junction error is ±1.5°C..."
question = "What is the maximum cold junction error for the SM 1231 Thermocouple module?"
user_message = f"""Now for the real task, solve the task in question block based on the context in context block.
Generate only the solution, do not generate anything else
<context>{context}</context>
<question>{question}</question>"""
messages = [
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_message},
]
input_ids = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
with torch.no_grad():
output = model.generate(input_ids, max_new_tokens=256, temperature=0.6, top_p=0.9)
response = tokenizer.decode(output[0][input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
Model Details
- Architecture: LlamaForCausalLM (1.24B parameters)
- Base model: meta-llama/Llama-3.2-1B-Instruct
- Context length: 131,072 tokens
- Precision: bfloat16
- Training method: Distillation via Distil Labs
Intended Use
This model is intended to be used as part of a RAG pipeline over Siemens S7-1200 documentation. Provide relevant context passages from the manual alongside user questions. The model was not trained for general-purpose chat or tasks outside this documentation domain.
Licenses
- Model weights: Distil Labs R&D License (non-commercial prototyping and R&D)
- Base model: Llama 3.2 Community License
- Teacher model: Apache 2.0
- Downloads last month
- 59
docker model run hf.co/distil-labs/distil-siemens-s7-1200-docs-llama-1b