Text Generation
PEFT
Safetensors
English
lora
qlora
sft
trl
text-classification
cve
cwe
vulnerability
security
conversational
Eval Results (legacy)
Instructions to use SASVAAI/Qwen-3.5-9B-CVE-to-CWE with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use SASVAAI/Qwen-3.5-9B-CVE-to-CWE with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "SASVAAI/Qwen-3.5-9B-CVE-to-CWE") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from SASVAAI/Qwen-3.5-9B-CVE-to-CWE: direct link, hf CLI and curl.
- Browser
- Download file 28.2 kB
-
https://huggingface.co/SASVAAI/Qwen-3.5-9B-CVE-to-CWE/resolve/main/README.md
- Command line
-
hf download hf://SASVAAI/Qwen-3.5-9B-CVE-to-CWE/README.md
-
curl -L -o README.md https://huggingface.co/SASVAAI/Qwen-3.5-9B-CVE-to-CWE/resolve/main/README.md
28.2 kB
| # ---- Identity ------------------------------------------------------------- | |
| base_model: Qwen/Qwen3.5-9B | |
| base_model_relation: adapter | |
| library_name: peft | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| license: apache-2.0 # verified: inherited from Qwen/Qwen3.5-9B, whose Hub | |
| # metadata declares license:apache-2.0 and which ships an | |
| # Apache-2.0 LICENSE file (copied here verbatim). | |
| # ---- Discovery ------------------------------------------------------------ | |
| tags: | |
| - lora | |
| - qlora | |
| - sft | |
| - trl | |
| - peft | |
| - text-classification | |
| - cve | |
| - cwe | |
| - vulnerability | |
| - security | |
| datasets: | |
| - exploitintel/cve-cwe-consensus | |
| metrics: | |
| - f1 | |
| - accuracy | |
| # ---- Structured evaluation ------------------------------------------------ | |
| model-index: | |
| - name: qwen3.5-9b-cve-cwe-lora | |
| results: | |
| - task: | |
| type: text-classification | |
| name: CVE description to CWE weakness class (117-way, generated as text) | |
| dataset: | |
| type: exploitintel/cve-cwe-consensus | |
| name: exploitintel/cve-cwe-consensus, 300-row natural-distribution sample of the validation split | |
| split: validation | |
| revision: 606cef101302fc2e0f69fd298f0de28b20a2aacf | |
| metrics: | |
| - type: f1 | |
| name: Macro-F1 over the union of gold and predicted classes (81 classes) | |
| value: 0.473288 | |
| - type: f1 | |
| name: Macro-F1 over gold classes only (68 classes) | |
| value: 0.563769 | |
| - type: accuracy | |
| name: Accuracy (= micro-F1 for single-label output) | |
| value: 0.71 | |
| # Qwen3.5-9B CVE-to-CWE Classifier (LoRA) | |
| Given the free-text description of a published CVE, emits the single CWE | |
| identifier that best characterises the root-cause weakness, as one line of | |
| text (`CWE-79`, `CWE-787`, ...), over the 117-class label space of | |
| [exploitintel/cve-cwe-consensus](https://huggingface.co/datasets/exploitintel/cve-cwe-consensus). | |
| This is a **LoRA adapter for** | |
| [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), trained with | |
| **QLoRA (4-bit NF4 base, bf16 compute)** via | |
| [TRL](https://github.com/huggingface/trl) SFT. | |
| > **The headline number is one draw from a noisy measurement.** This checkpoint | |
| > scored macro-F1 0.473 on 300 held-out rows, but the identical configuration | |
| > re-run later in the same search scored 0.395. Read | |
| > [How these values were chosen](#how-these-values-were-chosen) and | |
| > [Evaluation](#evaluation) before quoting anything, and do not report 0.473 as | |
| > the model's expected performance. | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | Developed by | SASVA AI Model Cognition Labs (MCL) Team | | |
| | Base model | [`Qwen/Qwen3.5-9B`](https://huggingface.co/Qwen/Qwen3.5-9B) | | |
| | Base revision | `c202236235762e1c871ad0ccb60c8ee5ba337b9a` | | |
| | Base parameters | 9.65B (dense; 9,653,104,368 in the safetensors index, ~19.3 GB bf16) | | |
| | Architecture family | `qwen3_5` (`Qwen3_5ForConditionalGeneration`, text tower `qwen3_5_text`) | | |
| | Adaptation | LoRA (`r=128`, `alpha=256`, `dropout=0.15`, rsLoRA off, DoRA off) | | |
| | Trainable modules | `in_proj_qkv`, `in_proj_z`, `out_proj` (Gated DeltaNet layers); `q_proj`, `k_proj`, `v_proj`, `o_proj` (attention layers); `gate_proj`, `up_proj`, `down_proj` (every layer) | | |
| | Excluded modules | none (the base is text-only in practice; no adapter tensor touches the vision tower) | | |
| | Training method | `qlora` (`--load-in-4bit`, run 3 / trial 23) | | |
| | Refinement | none | | |
| | Precision | 4-bit NF4 base with double quantisation, bf16 compute; adapter stored in float32 | | |
| | Language | English | | |
| | License | Apache-2.0 (inherited from the base model) | | |
| Trainable parameters: **320,864,256** across 200 modules, 3.2975% of the | |
| 9,730,678,000 parameters PEFT counted with the adapter attached. The adapter | |
| file is 1,283,518,408 bytes (400 tensors, `lora_A` + `lora_B` per module, all | |
| float32). Every tensor sits under `base_model.model.model.language_model`. | |
| **One Qwen 3.5 structural fact shapes the module list.** Confirmed against the | |
| base model's `config.json`: the 32 layers follow a 3-linear / 1-full pattern | |
| (`full_attention_interval: 4`), so 24 layers are Gated DeltaNet linear | |
| attention and expose `in_proj_qkv` / `in_proj_z` / `out_proj`, while the 8 | |
| full-attention layers (3, 7, 11, ..., 31) are grouped-query attention with 16 | |
| heads over 4 KV heads and expose `q_proj` / `k_proj` / `v_proj` / `o_proj`. | |
| The adapter's tensor counts match exactly: 48 per linear-attention projection | |
| (24 layers x A/B), 16 per attention projection (8 layers x A/B), 64 per MLP | |
| projection (32 layers x A/B). There are no MoE experts in this base. | |
| ## Intended use | |
| **Direct use.** Map a CVE description to one CWE id, for triage and | |
| labelling pipelines that already consume CWE ids. The 117-class label space is | |
| the dataset's; ids outside it were never seen in training. | |
| The model was trained on a specific prompt shape and that shape is part of the | |
| contract: | |
| - System prompt (verbatim): *"You are a CVE-to-CWE classifier. Given a CVE | |
| vulnerability description, identify the single root-cause CWE weakness class | |
| that best characterizes the flaw. Output exactly one CWE identifier (e.g., | |
| CWE-79) on a single line with no explanation or additional text."* | |
| - User turn: the instruction *"Classify the following CVE description into | |
| exactly one CWE weakness class. Reply with the CWE ID only, for example | |
| CWE-79."*, a blank line, then the CVE description inside a bare ```` ``` ```` | |
| fence. This is the exact string the evaluator rendered | |
| (`instruction + "\n\n```\n" + description + "\n```"`). | |
| - Applied through the tokenizer's chat template (`chat_template.jinja`, | |
| shipped in this repo) with `add_generation_prompt=True` and | |
| `enable_thinking=False`. Do not concatenate strings by hand. | |
| - The output is one line, `CWE-<n>`. In evaluation all 300 generations were a | |
| single well-formed id; the first line of the output is the prediction. | |
| - Decode greedily (`do_sample=False`) with a small budget; 64 new tokens is | |
| what the metric was scored with. | |
| ## How to get started | |
| ```python | |
| import torch | |
| from peft import PeftModel | |
| from transformers import AutoModelForImageTextToText, AutoTokenizer | |
| BASE = "Qwen/Qwen3.5-9B" | |
| ADAPTER = "SASVAAI/Qwen-3.5-9B-CVE-to-CWE" | |
| # NOTE: AutoModelForImageTextToText, not AutoModelForCausalLM. Qwen 3.5's | |
| # CausalLM mapping raises AttributeError: 'Qwen3_5Config' object has no | |
| # attribute 'vocab_size' on the transformers build used here; the | |
| # image-text-to-text mapping resolves to Qwen3_5ForConditionalGeneration and | |
| # loads the text model correctly. Training used the same class. | |
| tokenizer = AutoTokenizer.from_pretrained(ADAPTER, trust_remote_code=True) | |
| model = AutoModelForImageTextToText.from_pretrained( | |
| BASE, dtype=torch.bfloat16, device_map="auto", trust_remote_code=True | |
| ) | |
| model = PeftModel.from_pretrained(model, ADAPTER) | |
| model.eval() | |
| # Verbatim from the training invocation. Do not paraphrase. | |
| SYSTEM = ( | |
| "You are a CVE-to-CWE classifier. Given a CVE vulnerability description, " | |
| "identify the single root-cause CWE weakness class that best characterizes " | |
| "the flaw. Output exactly one CWE identifier (e.g., CWE-79) on a single line " | |
| "with no explanation or additional text." | |
| ) | |
| INSTRUCTION = ( | |
| "Classify the following CVE description into exactly one CWE weakness class. " | |
| "Reply with the CWE ID only, for example CWE-79." | |
| ) | |
| description = ( | |
| "A stored cross-site scripting flaw in the FAQ page lets an attacker inject " | |
| "script that runs in other users' browsers." | |
| ) | |
| messages = [ | |
| {"role": "system", "content": SYSTEM}, | |
| {"role": "user", "content": f"{INSTRUCTION}\n\n```\n{description}\n```"}, | |
| ] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, add_generation_prompt=True, enable_thinking=False, | |
| return_tensors="pt", return_dict=True, | |
| ).to(model.device) | |
| out = model.generate(**inputs, max_new_tokens=64, do_sample=False) | |
| print(tokenizer.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True).strip()) | |
| # -> CWE-79 | |
| ``` | |
| The base model is ~19 GB in bfloat16; one 24 GB-class GPU is enough for | |
| inference with the adapter. Loading the base in 4-bit with `bitsandbytes` | |
| (`load_in_4bit=True`, `bnb_4bit_quant_type="nf4"`, | |
| `bnb_4bit_use_double_quant=True`, `bnb_4bit_compute_dtype=torch.bfloat16`) | |
| reproduces the training-time numerics and needs about 7 GB. | |
| > Decoding matters. The metric was scored greedily with `max_new_tokens=64` | |
| > through the chat template with thinking disabled. No sampling setting was | |
| > validated, and enabling thinking changes the prompt the model sees. | |
| ## Training details | |
| **Data.** 700 training rows and 300 validation rows built from | |
| [exploitintel/cve-cwe-consensus](https://huggingface.co/datasets/exploitintel/cve-cwe-consensus) | |
| at revision `606cef101302fc2e0f69fd298f0de28b20a2aacf` by a deterministic | |
| script (`autocatalyst.datagen.cve_cwe_consensus`, seed 0). No model generated | |
| any training content. | |
| Row selection, as recorded in the builder's manifest: | |
| - Source rows are `{"messages": [system, user, assistant]}`; the user turn is | |
| the CVE description, the assistant turn the label. | |
| - Rows whose label is not exactly one `CWE-<n>` id were dropped (about 9% of | |
| the dataset are comma-separated multi-label lists). Rows whose id is outside | |
| the dataset's 117-class set, or whose description is empty, were dropped. | |
| Duplicate descriptions within a split were collapsed to the first. | |
| - **Train**: from the dataset's `train` split (50,074 raw rows, 44,337 kept | |
| after dropping 4,390 multi-label or malformed and 1,347 duplicate rows), | |
| 700 rows were sampled with one row guaranteed per class and the remainder | |
| weighted by the square root of class frequency, so the head classes do not | |
| crowd out the tail that macro-F1 scores. All 117 classes are present. | |
| - **Validation**: from the dataset's own `validation` split (11,052 raw rows, | |
| 8,878 kept), 300 rows sampled at the natural class distribution. 68 classes | |
| appear; every one of them is in the training set; no description text is | |
| shared with the 700 training rows (0 overlaps measured). | |
| | | | | |
| |---|---| | |
| | Train samples | 700 rows / 117 classes | | |
| | Validation samples | 300 rows / 68 classes | | |
| | Text overlap | 0 rows | | |
| | Prompt format | chat template + system prompt + instruction/fenced-description user turn (see Intended use) | | |
| | Loss masking | answer tokens only; prompt tokens set to `-100` | | |
| | Truncation | prompt left-truncated to fit `max_seq_len` 2048; answer never truncated | | |
| The dataset's head is heavy: in the 300 validation rows CWE-79 appears 49 | |
| times, CWE-862 22, CWE-89 17, CWE-200 and CWE-22 13 each, and 68 classes share | |
| the rest. | |
| An LLM (Claude Opus 4.6, `claude-opus-4-6`, via an internal inference gateway) | |
| proposed the hyperparameters the search tried. It generated no training | |
| content and computed no metric. | |
| ### Method | |
| | | | | |
| |---|---| | |
| | SFT method | `qlora` | | |
| | Base quantisation during training | 4-bit NF4, double quantisation, bf16 compute (`bitsandbytes`) | | |
| | Refinement stage | none | | |
| | Attention implementation | `flash_attention_2` for the 8 GQA layers; the 24 Gated DeltaNet layers ran the PyTorch fallback because `flash-linear-attention` was not installed | | |
| | Auto class | `AutoModelForImageTextToText` (see How to get started) | | |
| | Hardware | 8x NVIDIA H100 80GB HBM3, `torchrun --nproc_per_node=8` | | |
| The project allowed two methods for this run (`bf16_lora`, `qlora`) and the | |
| search tried both; see the trial table below. | |
| No refinement stage ran; the published weights are the SFT adapter. | |
| ### Final hyperparameters | |
| | Hyperparameter | Value | Source | | |
| |---|---|---| | |
| | `learning_rate` | 0.0002 | `[TRAIN]` cmdline | | |
| | `lr_scheduler_type` | cosine | `[TRAIN]` cmdline | | |
| | `num_train_epochs` | 5 | `[TRAIN]` cmdline, `trainer_state.json` | | |
| | `per_device_train_batch_size` | 2 | `[TRAIN]` cmdline, `trainer_state.json` | | |
| | `gradient_accumulation_steps` | 4 | `[TRAIN]` cmdline | | |
| | `max_seq_length` | 2048 | `[TRAIN]` cmdline | | |
| | `warmup_ratio` | 0.05 | `[TRAIN]` cmdline | | |
| | `weight_decay` | 0.05 | `[TRAIN]` cmdline | | |
| | `loraplus_lr_ratio` | 2.0 (LoRA `B` matrices at 2x the learning rate) | `[TRAIN]` cmdline | | |
| | `lora_r` / `lora_alpha` / `lora_dropout` | 128 / 256 / 0.15 | `adapter_config.json` | | |
| | `use_rslora` / `use_dora` | `false` / `false` | `adapter_config.json` | | |
| | `target_modules` | the 10 listed in Model details | `adapter_config.json` | | |
| | `load_in_4bit` | `true` (NF4, double quant, bf16 compute) | `[TRAIN]` cmdline | | |
| **Effective batch size: 64** (`2 x 4 x 8`). Optimizer steps: 55 (11 per | |
| epoch). | |
| `neftune_noise_alpha` (0.0), `use_liger_kernel`, `use_sample_packing` and | |
| `lora_init` (`default`) were left at their no-op defaults. KD parameters are | |
| omitted deliberately: this is a `qlora` run, not a distillation run. | |
| > Provenance note: every value above was recovered from the platform database | |
| > (`runs`, `experiments`, `events` tables for run 3) and cross-checked against | |
| > the literal `[TRAIN]` command line recorded in the run log and against the | |
| > shipped `adapter_config.json`. | |
| ### How these values were chosen | |
| > These hyperparameters were selected by an automated search | |
| > (`autocatalyst.cli.run_autoresearch`): an agent proposes one change at a | |
| > time, runs train then eval, and keeps or discards on `f1_macro` (higher is | |
| > better). | |
| Run 3 ran **46 trials in 7 h 17 m** (2026-09-08 20:29 to 2026-09-09 03:46); | |
| 35 scored and 11 errored. This checkpoint is **trial 23**, the run's best. | |
| **Every trial is comparable.** The search's five "data strategies" all | |
| re-adapted the same fixed 700-row training file (the `train_v1` ... `train_v5` | |
| files below hold the same 700 rows) and every trial was scored on the same 300 | |
| validation rows, so the whole table is one comparison. | |
| | # | Method | rsLoRA | r | Dropout | LR | Epochs | Grad accum | Warmup | Weight decay | LoRA+ ratio | Data | f1_macro | Kept | | |
| |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | |
| | 1 | bf16_lora | no | 16 | 0.05 | 2e-4 | 2 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.271028 | yes | | |
| | 2 | qlora | no | 16 | 0.05 | 2e-4 | 2 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.323589 | yes | | |
| | 3 | bf16_lora | no | 16 | 0.05 | 2e-4 | 2 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.271161 | no | | |
| | 4 | qlora | no | 16 | 0.05 | 1e-4 | 2 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.224824 | no | | |
| | 5 | qlora | no | 16 | 0.05 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.328347 | yes | | |
| | 6 | qlora | no | 16 | 0.05 | 2e-4 | 8 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.315655 | no | | |
| | 7 | qlora | no | 32 | 0.05 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.339796 | yes | | |
| | 8 | qlora | no | 32 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.358649 | yes | | |
| | 9 | qlora | no | 32 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.379137 | yes | | |
| | 10 | qlora | no | 64 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.411651 | yes | | |
| | 11 | qlora | yes | 64 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | 0.385129 | no | | |
| | 12 | qlora | no | 64 | 0.05 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.405714 | no | | |
| | 13 | qlora | no | 64 | 0.15 | 2e-4 | 5 | 8 | 0.05 | 0.01 | 2.0 | train_v1 | 0.377237 | no | | |
| | 14 | qlora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.440878 | yes | | |
| | 15 | qlora | no | 128 | 0.15 | 2e-4 | 8 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.377763 | no | | |
| | 16 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | error | no | | |
| | 17 | qlora | no | 128 | 0.15 | 1.5e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.386976 | no | | |
| | 18 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v1 | error | no | | |
| | 19 | qlora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v1 | 0.396817 | no | | |
| | 20 | qlora | no | 128 | 0.15 | 2e-4 | 8 | 4 | 0.05 | 0.01 | 2.0 | train_v2 | 0.407741 | no | | |
| | 21 | bf16_lora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v2 | 0.403981 | no | | |
| | 22 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 1.0 | train_v2 | error | no | | |
| | **23** | **qlora** | **no** | **128** | **0.15** | **2e-4** | **5** | **4** | **0.05** | **0.05** | **2.0** | **train_v2** | **0.473288** | **yes** | | |
| | 24 | qlora | no | 128 | 0.10 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v2 | 0.444999 | no | | |
| | 25 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v2 | error | no | | |
| | 26 | qlora | no | 128 | 0.15 | 2.5e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v2 | 0.427152 | no | | |
| | 27 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v2 | error | no | | |
| | 28 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v2 | error | no | | |
| | 29 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v3 | error | no | | |
| | 30 | qlora | no | 128 | 0.15 | 2e-4 | 8 | 4 | 0.05 | 0.05 | 2.0 | train_v3 | 0.384342 | no | | |
| | 31 | bf16_lora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v3 | 0.403688 | no | | |
| | 32 | qlora | no | 128 | 0.05 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v3 | 0.415561 | no | | |
| | 33 | qlora | no | 128 | 0.15 | 1.75e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v3 | 0.435022 | no | | |
| | 34 | qlora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.03 | 0.05 | 2.0 | train_v3 | 0.416758 | no | | |
| | 35 | qlora | no | 128 | 0.15 | 2e-4 | 8 | 4 | 0.05 | 0.05 | 2.0 | train_v4 | 0.406259 | no | | |
| | 36 | bf16_lora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v4 | 0.408250 | no | | |
| | 37 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v4 | error | no | | |
| | 38 | qlora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v4 | 0.395110 | no | | |
| | 39 | qlora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.01 | 2.0 | train_v4 | 0.437202 | no | | |
| | 40 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v4 | error | no | | |
| | 41 | qlora | no | 128 | 0.15 | 2e-4 | 8 | 4 | 0.05 | 0.05 | 2.0 | train_v5 | 0.391961 | no | | |
| | 42 | bf16_lora | no | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v5 | 0.428970 | no | | |
| | 43 | qlora | no | 128 | 0.05 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 2.0 | train_v5 | 0.423615 | no | | |
| | 44 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v5 | error | no | | |
| | 45 | qlora | yes | 128 | 0.15 | 2e-4 | 5 | 4 | 0.05 | 0.05 | 1.0 | train_v5 | error | no | | |
| | 46 | qlora | no | 128 | 0.15 | 2e-4 | 5 | 8 | 0.05 | 0.05 | 2.0 | train_v5 | 0.402578 | no | | |
| All trials: `lora_alpha = 2 x r`, cosine schedule, batch 2 per device, | |
| `max_seq_len` 2048. | |
| **What the search actually established.** | |
| - **Rank dominates, up to a point.** Holding QLoRA, 5 epochs, dropout 0.15, | |
| LR 2e-4, weight decay 0.01, LoRA+ 2.0: `r` 32 → 64 → 128 moved macro-F1 | |
| 0.379137 → 0.411651 → 0.440878 (#9, #10, #14). | |
| - **LoRA+ helped.** The one clean A/B at `r=32` (#8 → #9) moved 0.358649 → | |
| 0.379137 by putting the `B` matrices at 2x the learning rate. | |
| - **More epochs hurt.** 8 epochs lost to 5 every time it was tried at `r=128` | |
| (#15, #20, #30, #35, #41), by 0.01 to 0.06. | |
| - **4-bit versus bf16 base: no measurable difference.** At `r=16` the 4-bit | |
| base beat bf16 (#2 at 0.323589 against #1 and #3 at 0.271); at `r=128` with | |
| weight decay 0.05 the three bf16 runs (#31, #36, #42: 0.404 to 0.429) sit | |
| inside the two 4-bit runs of the same configuration (#38 at 0.395, #23 at | |
| 0.473). Quantising the base cost nothing this search could detect. | |
| - **rsLoRA at `r=128` diverges.** All 11 rsLoRA trials at `r=128` (#16, #18, | |
| #22, #25, #27, #28, #29, #37, #40, #44, #45) were aborted by the training | |
| watchdog at the third logging step with gradient norm above 10 (the | |
| configured limit; observed 12.3 in the first). The likely cause is the | |
| scaling rule: rsLoRA scales updates by `alpha / sqrt(r)` instead of | |
| `alpha / r`, which with `alpha = 2r` is `2 sqrt(r)`, about 22.6 at `r=128` | |
| against 2.0 without it, an 11x larger effective update. The one rsLoRA | |
| trial that finished (#11, `r=64`) scored below its non-rsLoRA sibling (#10). | |
| The 11 aborted trials are why the search shows 46 trials but 35 scores. | |
| **What it did not establish: the winning margin.** The search's decision | |
| "weight decay 0.01 → 0.05 improved 0.440878 → 0.473288" (#14 → #23) does not | |
| survive the re-runs. Identical configurations were run more than once because | |
| each "data strategy" restarted the inner loop on the same data: | |
| | Configuration (all QLoRA, r=128, dropout 0.15, LR 2e-4, 5 epochs, grad accum 4, warmup 0.05, LoRA+ 2.0) | Trials | f1_macro | Spread | | |
| |---|---|---|---| | |
| | weight decay 0.05 (**this checkpoint's config**) | #23, #38 | 0.473288, 0.395110 | 0.078 | | |
| | weight decay 0.01 | #14, #19, #39 | 0.440878, 0.396817, 0.437202 | 0.044 | | |
| | weight decay 0.05, bf16 base instead of 4-bit | #31, #36, #42 | 0.403688, 0.408250, 0.428970 | 0.025 | | |
| | weight decay 0.05, 8 epochs | #30, #35, #41 | 0.384342, 0.406259, 0.391961 | 0.022 | | |
| | weight decay 0.05, dropout 0.05 | #32, #43 | 0.415561, 0.423615 | 0.008 | | |
| Seeds were not pinned, so these differ only in initialisation and data order. | |
| The 0.078 gap between #23 and #38 is larger than the 0.032 "improvement" the | |
| search kept, and larger than most differences in the table. Read 0.473 as the | |
| high draw of a configuration whose expected score is somewhere in the low-to-mid | |
| 0.4s, and treat any two trials within about 0.05 of each other as tied. | |
| **Search space.** Six knobs were varied (`LORA_R`, `LORA_DROPOUT`, | |
| `LEARNING_RATE`, `EPOCHS`, `WEIGHT_DECAY`, `LORAPLUS_LR_RATIO`) plus one | |
| `GRAD_ACCUM` probe (#13, #46), one warmup probe (#34), the bf16-vs-4-bit | |
| method switch, and the rsLoRA attempts. `LR_SCHEDULER`, `MAX_SEQ_LEN`, | |
| `BATCH_SIZE`, `NEFTUNE_NOISE_ALPHA`, `USE_DORA`, `LORA_INIT`, | |
| `USE_LIGER_KERNEL` and `USE_SAMPLE_PACKING` were never moved. | |
| **Observed training metrics** (this checkpoint). | |
| | | | | |
| |---|---| | |
| | Final train loss (mean over the run) | 0.6052066683769226 | | |
| | Final eval loss (teacher-forced, answer tokens) | 0.9717274904251099 | | |
| | Eval mean token accuracy | 0.8136 | | |
| | Train runtime | 387.7625 s | | |
| | Total FLOPs | 4.311845712166912e+16 | | |
| | Throughput | 9.026 samples/s, 0.142 steps/s | | |
| 55 optimizer steps ran. The logged train loss fell from 1.2962 at step 10 (grad | |
| norm 1.06) to the run mean of 0.6052; the reported train loss is the mean over | |
| the run, not a converged value. | |
| ## Evaluation | |
| **Protocol.** All 300 validation rows, greedy decoding, `max_new_tokens=64`, | |
| prompts rendered through the chat template with `enable_thinking=False`. The | |
| project's `classification` evaluator takes the first line of the generation as | |
| the predicted label and compares it to the gold id. Generation ran through | |
| Hugging Face `generate` in batches of 4 (`max_input_len` 4096) after vLLM | |
| 0.19.1 refused the LoRA on this architecture and the evaluator fell back; the | |
| whole pass took 63 s on 8 GPUs. Every one of the 300 outputs was a single | |
| well-formed `CWE-<n>` id, so no prediction was lost to formatting, and the | |
| 64-token budget cut nothing (the longest output is one id). | |
| | Metric | Value | | |
| |---|---| | |
| | Macro-F1, union of gold and predicted classes (**the search metric**) | 0.473288 | | |
| | Macro-F1, gold classes only | 0.563769 | | |
| | Accuracy (= micro-F1) | 0.71 (213 / 300) | | |
| | Classes in gold / predicted / union | 68 / 72 / 81 | | |
| **Two macro-F1 numbers, one convention.** The evaluator averages F1 over the | |
| *union* of gold and predicted classes, so the 13 classes the model predicted | |
| that never occur in the 300 gold rows each contribute an F1 of 0 and pull the | |
| macro average down from 0.5638 to 0.4733. Both are reported; the union | |
| convention is the one the search optimised and the one the published baselines | |
| below appear to use, but check before comparing. | |
| **Published comparison points.** The dataset authors report, on their own | |
| evaluation of the same dataset (a much larger split than these 300 rows): | |
| | Model | Method | Micro-F1 | Macro-F1 | | |
| |---|---|---|---| | |
| | Qwen3-32B ([exploitintel/cve-cwe-qwen3-32b](https://huggingface.co/exploitintel/cve-cwe-qwen3-32b)) | QLoRA r=16 | 0.729 | 0.595 | | |
| | Qwen3-8B (same author, same recipe) | QLoRA r=16 | 0.702 | 0.511 | | |
| | **This adapter** (Qwen3.5-9B, QLoRA r=128, 700 training rows) | | **0.71** | **0.473** | | |
| Those models trained on the full ~44K-row training split; this adapter saw | |
| 700 rows. The micro-F1 is in the same range; the macro-F1 is 0.04 to 0.12 | |
| lower, which is the long tail this adapter had one to twelve examples of per | |
| class to learn from. The evaluation sets also differ in size and composition, | |
| so read this as context, not a controlled comparison. | |
| **Sibling run on the same 700 / 300 split.** A Gemma 4 E4B adapter trained by | |
| the same platform on this split reached accuracy 0.697 and union macro-F1 | |
| 0.4587, essentially tied with this checkpoint given the spread above. | |
| **Baseline for comparison. Not measured.** The untuned `Qwen/Qwen3.5-9B` was | |
| never scored on these 300 rows, so nothing here quantifies how much of the | |
| score the fine-tuning is responsible for. This is the most important gap in | |
| this card. | |
| **This is a validation split the search selected against.** 35 trials were | |
| scored on these same 300 rows and the best was kept, so expect optimistic bias | |
| on top of the re-run spread already described. The rows are drawn from the | |
| dataset authors' own validation split, so they are unseen CVEs from the same | |
| period as training, not future CVEs. | |
| **The evaluation set is reproducible.** `predictions.jsonl` in this repo holds | |
| every one of the 300 rows: instruction, description, prediction and gold. The | |
| builder's manifest (dataset revision, seed, drop counts, per-class counts) is | |
| summarised under Training details. | |
| ## Limitations and bias | |
| **One number, wide error bars.** The same configuration scored 0.473 and 0.395 | |
| in two runs. Anyone deploying this should re-evaluate on their own data rather | |
| than trust either figure. | |
| **No baseline, so no established gain.** See Evaluation. | |
| **Head classes dominate what accuracy measures.** CWE-79 alone is 16% of the | |
| validation rows. A model that got only the top ten classes right would post a | |
| respectable accuracy and a poor macro-F1; the two numbers here disagree by | |
| 0.24 for that reason. | |
| **Tail classes were barely trained.** 117 classes over 700 rows means many | |
| classes had a single training example. Expect the model to fall back to a | |
| frequent neighbour (CWE-20 for input validation issues, CWE-200 for disclosure) | |
| when the description is ambiguous, and to emit ids it saw rarely with low | |
| reliability. It also produced 13 ids in evaluation that never occur in the 300 | |
| gold rows; some may be reasonable alternative labels, some are wrong. | |
| **Only one label.** Real CVEs often carry two CWE ids (about 9% of the source | |
| dataset). The training data dropped those rows, so the model always commits to | |
| one. | |
| **Prompt shape is the contract.** Change the system prompt, the instruction | |
| sentence, the code fence, or enable thinking, and you are evaluating a model | |
| nobody measured. | |
| **Domain narrowness.** English CVE descriptions in NVD / CNA style. Advisories | |
| in other formats, other languages, source-code inputs and exploit write-ups are | |
| unmeasured. | |
| **Inherits all biases and limitations of the base model.** This adapter changes | |
| 3.3% of the parameters and was not evaluated for safety or fairness. The base | |
| model's own card governs those properties. | |
| ## Environmental impact | |
| | | | | |
| |---|---| | |
| | Hardware | 8x NVIDIA H100 80GB HBM3 | | |
| | Training time | 6.46 minutes (387.7625 s) | | |
| | Cloud provider / region | on-premise | | |
| Covers this trial only. The full 46-trial search that selected it took 7 h 17 m | |
| on the same hardware. | |
| ## Framework versions | |
| - PEFT 0.18.1 | |
| - TRL: 1.0.0 | |
| - Transformers: 5.7.0.dev0 (git main) | |
| - Pytorch: 2.5.1+cu121 | |
| - bitsandbytes: 0.49.2 | |
| - flash-attn: 2.8.3 | |
| - Python: 3.12 | |
| PEFT's version is the one recorded in `adapter_config.json` at save time; the | |
| rest are the pinned versions of the training environment. `transformers` is a | |
| git-main build: the `qwen3_5` architecture is not in the stable PyPI release. | |
| ## Citation | |
| ```bibtex | |
| @misc{qwen35_cve_cwe_lora_2026, | |
| title = {Qwen3.5-9B CVE-to-CWE Classifier LoRA}, | |
| author = {Banerjee, Aaron and Anbuselvan, Pooja and Jodhpurkar, Om}, | |
| year = {2026}, | |
| url = {https://huggingface.co/SASVAAI/Qwen-3.5-9B-CVE-to-CWE} | |
| } | |
| ``` | |