kaleidoscopicwhether commited on
Commit
038c760
·
verified ·
1 Parent(s): 12431e1

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +51 -44
README.md CHANGED
@@ -1,52 +1,59 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  # SVG Generation Model Weights
2
 
3
  Model weights for the DL Spring 2026 Kaggle Competition — Text-to-SVG Generation.
4
 
5
- **Team:** Ivan Aristy (NYU Tandon)
6
- **Final Score:** 16.87/100 (6th place / 58 teams)
7
- **Base Model:** Qwen/Qwen2.5-Coder-1.5B-Instruct
 
8
 
9
  ## Models
10
 
11
- | Model | Type | Kaggle Score | Description |
12
  |---|---|---|---|
13
- | `componly-r32-adapter` | LoRA adapter | **16.87** | **Best model.** Load on `merged-1.5b-r16`. |
14
- | `refined-7000` | Full model | 16.26 | Full fine-tune, loss 0.308 |
15
- | `r16-3epoch` | LoRA adapter | 15.47 | First adapter, load on Qwen2.5-Coder-1.5B |
16
- | `mixed-r32-adapter` | LoRA adapter | 14.64 | Mixed data experiment |
17
- | `codegen-1.5b` | Full model | 12.26 | Code generation experiment |
18
- | `merged-1.5b-r16` | Full model | — | Base model with r16 knowledge baked in |
19
-
20
- ## Usage (Best Model)
21
-
22
- ```python
23
- from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
24
- from peft import PeftModel
25
- import torch
26
-
27
- bnb_config = BitsAndBytesConfig(
28
- load_in_4bit=True, bnb_4bit_quant_type="nf4",
29
- bnb_4bit_compute_dtype=torch.bfloat16, bnb_4bit_use_double_quant=True,
30
- )
31
-
32
- # Load merged base
33
- tokenizer = AutoTokenizer.from_pretrained("kaleidoscopicwhether/svg-gen-weights", subfolder="merged-1.5b-r16")
34
- model = AutoModelForCausalLM.from_pretrained(
35
- "kaleidoscopicwhether/svg-gen-weights", subfolder="merged-1.5b-r16",
36
- quantization_config=bnb_config, device_map="auto", torch_dtype=torch.bfloat16,
37
- )
38
-
39
- # Load best adapter
40
- model = PeftModel.from_pretrained(model, "kaleidoscopicwhether/svg-gen-weights", subfolder="componly-r32-adapter")
41
- model.eval()
42
-
43
- # Generate
44
- prompt = "<|im_start|>system\nOutput valid SVG code only.<|im_end|>\n<|im_start|>user\nA red circle<|im_end|>\n<|im_start|>assistant\n"
45
- inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
46
- output = model.generate(**inputs, max_new_tokens=1024, do_sample=False, repetition_penalty=1.1)
47
- print(tokenizer.decode(output[0], skip_special_tokens=True))
48
- ```
49
-
50
- ## Training Details
51
-
52
- See [GitHub repo](https://github.com/ivanearisty/svg-gen) and paper in `report/main.pdf`.
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ tags:
6
+ - svg
7
+ - text-to-svg
8
+ - qwen2
9
+ - qlora
10
+ - lora
11
+ - fine-tuning
12
+ - code-generation
13
+ base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
14
+ pipeline_tag: text-generation
15
+ library_name: transformers
16
+ ---
17
+
18
  # SVG Generation Model Weights
19
 
20
  Model weights for the DL Spring 2026 Kaggle Competition — Text-to-SVG Generation.
21
 
22
+ **Author:** Ivan Aristy (NYU Tandon, CS-GY 9223 / ECE-GY 7123)
23
+ **Base Model:** [Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct)
24
+ **Best Public Score:** 16.87 / 100
25
+ **GitHub:** [ivanearisty/svg-gen](https://github.com/ivanearisty/svg-gen)
26
 
27
  ## Models
28
 
29
+ | Model | Type | Public Score | Description |
30
  |---|---|---|---|
31
+ | [componly-r32-adapter](./componly-r32-adapter) | LoRA adapter | **16.87** | **Best model.** Requires `merged-1.5b-r16` as base. LoRA r=32, alpha=64, trained on 45k competition-only samples. |
32
+ | [merged-1.5b-r16](./merged-1.5b-r16) | Full model | — | Qwen2.5-Coder-1.5B with Round 1 LoRA r=16 adapter permanently merged into weights. Base for the best adapter. |
33
+ | [refined-7000](./refined-7000) | Full model | 16.26 | Full fine-tune from merged base, checkpoint 7000, CE loss 0.308. Standalone model. |
34
+ | [r16-3epoch](./r16-3epoch) | LoRA adapter | 15.47 | Round 1 adapter. LoRA r=16, 3 epochs on 46k competition data. Load on `Qwen/Qwen2.5-Coder-1.5B-Instruct`. |
35
+ | [mixed-r32-adapter](./mixed-r32-adapter) | LoRA adapter | 14.64 | LoRA r=32, trained on 76k mixed data (competition + external). External data hurt performance. |
36
+ | [codegen-1.5b](./codegen-1.5b) | Full model | 12.26 | Code generation experiment — model outputs Python code instead of raw SVG. |
37
+
38
+ ## Training Strategy
39
+
40
+ We use an iterative **merge-and-retrain** approach:
41
+
42
+ 1. **Round 1:** LoRA r=16 on Qwen2.5-Coder-1.5B (46k competition data, 3 epochs)
43
+ 2. **Merge** adapter into base weights
44
+ 3. **Round 2:** Fresh LoRA r=32 on merged base (45k competition data, 2 epochs) — **best model**
45
+ 4. **Round 3:** Merge Round 2 adapter, full fine-tune on DGX Spark
46
+
47
+ ## Key Findings
48
+
49
+ - **System prompt is critical:** Removing it drops score from 53.8 to 18.8 (local ablation)
50
+ - **Less is more:** Minimal system prompt ("Output valid SVG code only.") outperforms verbose prompts
51
+ - **LoRA > Full fine-tune:** Despite lower CE loss, full fine-tuning scores worse than LoRA
52
+ - **Competition data only:** External datasets (SVGX, OmniSVG) degraded performance
53
+ - **Greedy decoding optimal:** Any sampling or penalty variation hurts structured SVG output
54
+
55
+ ## Hardware
56
+
57
+ - **Training (QLoRA):** NVIDIA RTX 2000 Ada (16GB VRAM)
58
+ - **Training (Full FT):** NVIDIA DGX Spark (128GB unified memory)
59
+ - **Inference:** RTX 2000 Ada, 4-bit quantized, ~20s per SVG