Instructions to use jgalego/reqlint-smollm3-3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jgalego/reqlint-smollm3-3b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jgalego/reqlint-smollm3-3b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jgalego/reqlint-smollm3-3b") model = AutoModelForCausalLM.from_pretrained("jgalego/reqlint-smollm3-3b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jgalego/reqlint-smollm3-3b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jgalego/reqlint-smollm3-3b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgalego/reqlint-smollm3-3b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jgalego/reqlint-smollm3-3b
- SGLang
How to use jgalego/reqlint-smollm3-3b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jgalego/reqlint-smollm3-3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgalego/reqlint-smollm3-3b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jgalego/reqlint-smollm3-3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jgalego/reqlint-smollm3-3b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jgalego/reqlint-smollm3-3b with Docker Model Runner:
docker model run hf.co/jgalego/reqlint-smollm3-3b
reqlint-smollm3-3b
HuggingFaceTB/SmolLM3-3B fine-tuned with LoRA to check a requirement for common writing defects and rewrite it in EARS form. Where the rewrite needs information the original does not give, such as a time limit or the system that acts, it leaves a <placeholder> instead of inventing one.
Review aid only. Not a substitute for requirements review under DO-178C, ISO 26262, EN 50128, IEC 61508 or any other process. A clean result does not mean a requirement is correct, complete or verifiable.
🚀 Usage
import json
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "jgalego/reqlint-smollm3-3b"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="auto")
messages = [
{"role": "system", "content": "Check the requirement for these defects: compound, escape, no_unit, passive, pronoun, tbd, vague, weak_verb. Answer with JSON: {\"defects\": [...], \"rewrite\": [...]}. The rewrite lists one EARS requirement per line, with a \u003cplaceholder\u003e wherever information is missing. Leave both lists empty when there are no defects."},
{"role": "user", "content": "The battery management system shall open the main contactor within 2 s and measure the cell voltages every 1."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
)
out = model.generate(**inputs, max_new_tokens=200, do_sample=False)
print(json.loads(tokenizer.decode(out[0, inputs["input_ids"].shape[1] :], skip_special_tokens=True)))
Expected output:
{"defects": ["compound", "no_unit"], "rewrite": ["The battery management system shall open the main contactor within 2 s.", "The battery management system shall measure the cell voltages every 1 <unit>."]}
🔎 Defects
| Defect | Meaning |
|---|---|
compound |
More than one requirement in one statement |
escape |
A clause that lets the requirement be skipped: "if possible", "where practical" |
no_unit |
A number with no unit |
passive |
Passive voice with no actor: "the alarm shall be raised" |
pronoun |
A pronoun whose referent is unclear |
tbd |
A placeholder such as TBD or TBC |
vague |
An unmeasurable word where a value is needed: "quickly", "accurately" |
weak_verb |
should, may, can, will and similar instead of shall |
🏋️ Training
| Base | HuggingFaceTB/SmolLM3-3B |
| Data | 20,000 synthetic requirements |
| Method | LoRA, r=16, alpha=32, all linear layers, merged |
| Steps | 1250 (1.0 epochs), batch size 16 |
| Learning rate | 0.0002, cosine |
| Mean train loss | 0.0096 |
| Hardware | NVIDIA A10G, 47 min |
The generator writes requirements in the six EARS patterns for rail, automotive, aviation, space and energy systems, then injects up to two defects into each. Labels and rewrites come from the same structure, so they are exact.
📊 Results
Defect detection, micro-F1 over the 8 classes. rules is a keyword and regex linter built from the word lists the generator uses for training.
- Real: hand-labelled requirements from public-domain and openly licensed documents, in jgalego/reqlint-real.
- Synthetic, unseen: maritime and medical systems, with vague words, escape clauses, weak verbs and placeholders that never appear in training.
- Synthetic, seen: same domains and word lists as training, different requirements.
| Model | Real | Synthetic, unseen | Synthetic, seen | Rewrite exact, unseen |
|---|---|---|---|---|
| jgalego/reqlint-smollm3-3b | 0.57 | 0.921 | 1.0 | 0.93 |
| rules | 0.44 | 0.602 | 0.996 | 0.0 |
| HuggingFaceTB/SmolLM3-3B | 0.0 | 0.0 | 0.0 | 0.0 |
Per-defect F1 on real requirements:
| Model | compound | escape | no_unit | passive | pronoun | tbd | vague | weak_verb |
|---|---|---|---|---|---|---|---|---|
| jgalego/reqlint-smollm3-3b | 0.63 | 0.513 | 0.286 | 0.63 | 0.6 | 0.0 | 0.505 | 0.625 |
| rules | 0.222 | 0.125 | 0.182 | 0.866 | 0.222 | 0.0 | 0.162 | 0.483 |
| HuggingFaceTB/SmolLM3-3B | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
⚠️ Limitations
- One requirement at a time, in English. It does not check consistency across a specification.
- It flags how a requirement is written, not whether it is technically right.
- The training data is synthetic. The real set is small, so treat its scores as indicative.
- Downloads last month
- 240
Model tree for jgalego/reqlint-smollm3-3b
Dataset used to train jgalego/reqlint-smollm3-3b
Collection including jgalego/reqlint-smollm3-3b
Evaluation results
- Defect micro-F1 on Real requirementsself-reported0.570
- Defect micro-F1 on Synthetic, unseenself-reported0.921