krzysztofwos commited on
Commit
582b8cc
·
verified ·
1 Parent(s): 5f4dd52

Replace placeholder model card with proper documentation

Browse files
Files changed (1) hide show
  1. README.md +39 -52
README.md CHANGED
@@ -19,7 +19,7 @@ A LoRA fine-tuned adapter for [LiquidAI/LFM2.5-1.2B-Instruct](https://huggingfac
19
 
20
  ## Model Description
21
 
22
- This adapter teaches LFM2.5-1.2B-Instruct to respond in the structured Thought + Code format required by smolagents CodeAgent:
23
 
24
  ````
25
  Thought: I need to calculate this.
@@ -31,66 +31,53 @@ final_answer(result)
31
 
32
  ### Key Features
33
 
34
- - **Base Model**: LiquidAI/LFM2.5-1.2B-Instruct (1.2B parameter)
35
- - **Format Compliance**: N/A with minimal prompt
36
- - **Answer Accuracy**: N/A on evaluation tasks
37
- - **Adapter Size**: ~47MB (LoRA rank=16, alpha=32)
38
 
39
  ## Training Details
40
 
41
  ### Training Data
42
 
43
- - **81 successful CodeAgent trajectories** generated using Claude 3 Haiku as the teacher model
44
- - Tasks include mathematical reasoning, string manipulation, and general problem-solving
45
  - Each trajectory demonstrates the Thought → Code → Observation → final_answer pattern
 
46
 
47
  ### Training Configuration
48
 
49
- | Parameter | Value |
50
- | -------------------- | ------------------ |
51
- | LoRA Rank | 16 |
52
- | LoRA Alpha | 32 |
53
- | Target Modules | w1, w2, w3, q_proj, k_proj, v_proj, out_proj, in_proj |
54
- | Trainable Parameters | ~25M |
55
- | Training Steps | 3 epochs |
56
- | Learning Rate | 1e-4 |
57
- | Batch Size | 2 (grad accum 4) |
58
- | Max Sequence Length | 8192 |
59
- | Hardware | GPU |
60
- | Training Time | see ablation report |
61
 
62
  ### Training Framework
63
 
64
  - [TRL](https://github.com/huggingface/trl) SFTTrainer
65
  - [PEFT](https://github.com/huggingface/peft) for LoRA
66
 
67
- ## Evaluation Results
68
 
69
- ### Prompt Mode Comparison
70
 
71
- | Prompt Mode | Format Compliance | Answer Accuracy |
72
- | ----------- | --------------------------- | ------------------------- |
73
- | **Minimal** | N/A | N/A |
74
- | Default | N/A | N/A |
75
- | None | N/A | N/A |
 
 
76
 
77
- The model performs best with the **minimal prompt** (~95 tokens), demonstrating successful prompt distillation.
78
-
79
- ### Minimal Prompt Template
80
-
81
- ````text
82
- You are a CodeAgent that solves tasks by writing and executing Python code.
83
-
84
- Always respond with Thought + Python code block. Example:
85
-
86
- Thought: I need to calculate this.
87
- ```python
88
- result = 2 + 2
89
- final_answer(result)
90
- ```
91
-
92
- Call final_answer(result) when done. Now Begin!
93
- ````
94
 
95
  ## Usage
96
 
@@ -119,8 +106,8 @@ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
119
  outputs = model.generate(
120
  **inputs,
121
  max_new_tokens=512,
122
- temperature=0.3,
123
- min_p=0.15,
124
  )
125
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
126
  ```
@@ -146,23 +133,23 @@ print(result)
146
 
147
  ## Intended Use
148
 
149
- - Code-assisted problem solving
150
  - Mathematical reasoning tasks
151
  - Automated code generation following structured formats
152
- - Research into prompt distillation and small model fine-tuning
153
 
154
  ## Limitations
155
 
156
- - **Requires specific prompt format**: Works best with minimal prompt template
157
- - **Limited reasoning depth**: 1.2B parameter model has constrained reasoning capabilities compared to larger models
158
  - **English only**: Trained on English-language tasks
 
159
 
160
  ## Citation
161
 
162
  If you use this model, please cite:
163
 
164
  ```bibtex
165
- @misc{lfm25_1.2b_codeagent_haiku_default,
166
  author = {krzysztofwos},
167
  title = {LFM25-1.2B-CodeAgent-haiku-default},
168
  year = {2025},
@@ -173,6 +160,6 @@ If you use this model, please cite:
173
 
174
  ## Acknowledgments
175
 
176
- - [LiquidAI](https://www.liquid.ai/) for the LFM2.5-1.2B-Instruct base model
 
177
  - [Hugging Face](https://huggingface.co/) for smolagents, TRL, and PEFT
178
- - Training performed as part of CodeAgent prompt distillation research
 
19
 
20
  ## Model Description
21
 
22
+ This adapter teaches LFM2.5-1.2B to respond in the structured Thought + Code format required by smolagents CodeAgent:
23
 
24
  ````
25
  Thought: I need to calculate this.
 
31
 
32
  ### Key Features
33
 
34
+ - **Base Model**: LiquidAI/LFM2.5-1.2B-Instruct (1.17B parameters)
35
+ - **Token Accuracy**: 73.8% on training data
36
+ - **Single-Turn Rate**: 55.6% of tasks solved in one turn
37
+ - **Adapter Size**: ~42MB (LoRA rank=16, alpha=32)
38
 
39
  ## Training Details
40
 
41
  ### Training Data
42
 
43
+ - **81 successful CodeAgent trajectories** generated using Claude 3 Haiku
44
+ - Tasks include mathematical reasoning, string manipulation, file operations, and general problem-solving
45
  - Each trajectory demonstrates the Thought → Code → Observation → final_answer pattern
46
+ - Average trajectory length: 2957 tokens
47
 
48
  ### Training Configuration
49
 
50
+ | Parameter | Value |
51
+ | -------------------- | ----------------------------------------------------- |
52
+ | LoRA Rank | 16 |
53
+ | LoRA Alpha | 32 |
54
+ | Target Modules | w1, w2, w3, q_proj, k_proj, v_proj, out_proj, in_proj |
55
+ | Trainable Parameters | 11.1M (0.94% of base model) |
56
+ | Epochs | 3 |
57
+ | Learning Rate | 1e-4 |
58
+ | Batch Size | 2 (effective 8 with gradient accumulation) |
59
+ | Max Sequence Length | 8192 |
60
+ | Hardware | NVIDIA RTX 3090 (24GB) |
61
+ | Training Time | 461s |
62
 
63
  ### Training Framework
64
 
65
  - [TRL](https://github.com/huggingface/trl) SFTTrainer
66
  - [PEFT](https://github.com/huggingface/peft) for LoRA
67
 
68
+ ## Ablation Study Results
69
 
70
+ This model is part of a teacher ablation study comparing different Claude models and prompting strategies:
71
 
72
+ | Teacher Config | Token Accuracy | Training Time | Avg Trajectory Tokens |
73
+ | --------------- | -------------- | ------------- | --------------------- |
74
+ | haiku-default | 73.8% | 461s | 2,957 |
75
+ | sonnet-default | 90.6% | 260s | 3,054 |
76
+ | opus4-default | 90.2% | 260s | 2,971 |
77
+ | sonnet4-terse | 94.9% | 93s | 632 |
78
+ | **opus4-terse** | **95.0%** | **76s** | **613** |
79
 
80
+ **Key Finding**: Terse, focused trajectories from capable teachers transfer significantly better to small models.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
81
 
82
  ## Usage
83
 
 
106
  outputs = model.generate(
107
  **inputs,
108
  max_new_tokens=512,
109
+ temperature=0.1,
110
+ top_p=0.1,
111
  )
112
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
113
  ```
 
133
 
134
  ## Intended Use
135
 
136
+ - Code-assisted problem solving with small, efficient models
137
  - Mathematical reasoning tasks
138
  - Automated code generation following structured formats
139
+ - Research into prompt distillation and teacher model selection
140
 
141
  ## Limitations
142
 
143
+ - **1.2B model constraints**: Limited reasoning depth compared to larger models
 
144
  - **English only**: Trained on English-language tasks
145
+ - **CodeAgent format specific**: Optimized for smolagents Thought/Code/final_answer pattern
146
 
147
  ## Citation
148
 
149
  If you use this model, please cite:
150
 
151
  ```bibtex
152
+ @misc{lfm25-codeagent-haiku_default-2025,
153
  author = {krzysztofwos},
154
  title = {LFM25-1.2B-CodeAgent-haiku-default},
155
  year = {2025},
 
160
 
161
  ## Acknowledgments
162
 
163
+ - [LiquidAI](https://www.liquid.ai/) for the LFM2.5 base model
164
+ - [Anthropic](https://www.anthropic.com/) for Claude 3 Haiku (teacher model)
165
  - [Hugging Face](https://huggingface.co/) for smolagents, TRL, and PEFT