Instructions to use MSweetbread/qwen2.5-1.5b-chess-qlora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MSweetbread/qwen2.5-1.5b-chess-qlora with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("MSweetbread/qwen2.5-1.5b-chess-qlora", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
---
|
| 2 |
library_name: transformers
|
| 3 |
-
tags: [chess, fine-tuned,
|
| 4 |
base_model: Qwen/Qwen2.5-1.5B
|
| 5 |
---
|
| 6 |
|
|
@@ -14,14 +14,14 @@ For this assignment, only a free tier of Google Colab was allowed to train your
|
|
| 14 |
Additionally, the use of a pre-trained chess LLM was not permitted.
|
| 15 |
|
| 16 |
## Model Summary
|
| 17 |
-
A
|
| 18 |
Given a board state in the Forsyth-Edwards Notation (FEN) notation, the model outputs a move in the Universal Chess Interface (UCI) format.
|
| 19 |
|
| 20 |
## Model Architecture:
|
| 21 |
- **Base model:** Qwen2.5-1.5B (Causal LM)
|
| 22 |
- **Architecture:** Transformer with RoPE, SwiGLU, RMSNorm, Attention QKV bias and tied word embeddings
|
| 23 |
- **Total parameters:** 1.54B (1.31B non-embedding)
|
| 24 |
-
- **Trainable
|
| 25 |
- **Layers:** 28 | **Context length:** 32,768 tokens
|
| 26 |
|
| 27 |
## Training Data
|
|
@@ -36,7 +36,7 @@ Lichess entries were filtered to engine depth ≥ 16 to ensure high-quality
|
|
| 36 |
move annotations, at the cost of reduced dataset size.
|
| 37 |
|
| 38 |
## Training Procedure
|
| 39 |
-
- **Method:**
|
| 40 |
- **Training samples:** 227,134
|
| 41 |
- **Effective batch size:** 64
|
| 42 |
- **Training steps:** 7,098 (2 epochs)
|
|
@@ -61,7 +61,7 @@ print(tokenizer.decode(output[0], skip_special_tokens=True))
|
|
| 61 |
## Limitations
|
| 62 |
- Training loss of 1.089 suggests the model is not highly confident in its
|
| 63 |
predictions and will produce suboptimal moves in many positions.
|
| 64 |
-
-
|
| 65 |
updated; chess knowledge is partially constrained by the base LLM's pretraining.
|
| 66 |
- Training data consists primarily of high-level Lichess games (depth ≥ 16),
|
| 67 |
meaning the model is tuned on strong engine moves and may struggle with
|
|
|
|
| 1 |
---
|
| 2 |
library_name: transformers
|
| 3 |
+
tags: [chess, fine-tuned, qlora, fen, uci]
|
| 4 |
base_model: Qwen/Qwen2.5-1.5B
|
| 5 |
---
|
| 6 |
|
|
|
|
| 14 |
Additionally, the use of a pre-trained chess LLM was not permitted.
|
| 15 |
|
| 16 |
## Model Summary
|
| 17 |
+
A QLoRA fine-tuned version of the Qwen2.5-1.5B LLM for chess move prediction.
|
| 18 |
Given a board state in the Forsyth-Edwards Notation (FEN) notation, the model outputs a move in the Universal Chess Interface (UCI) format.
|
| 19 |
|
| 20 |
## Model Architecture:
|
| 21 |
- **Base model:** Qwen2.5-1.5B (Causal LM)
|
| 22 |
- **Architecture:** Transformer with RoPE, SwiGLU, RMSNorm, Attention QKV bias and tied word embeddings
|
| 23 |
- **Total parameters:** 1.54B (1.31B non-embedding)
|
| 24 |
+
- **Trainable QLoRA parameters:** 4,358,144 (0.49% of total)
|
| 25 |
- **Layers:** 28 | **Context length:** 32,768 tokens
|
| 26 |
|
| 27 |
## Training Data
|
|
|
|
| 36 |
move annotations, at the cost of reduced dataset size.
|
| 37 |
|
| 38 |
## Training Procedure
|
| 39 |
+
- **Method:** QLoRA fine-tuning (PEFT)
|
| 40 |
- **Training samples:** 227,134
|
| 41 |
- **Effective batch size:** 64
|
| 42 |
- **Training steps:** 7,098 (2 epochs)
|
|
|
|
| 61 |
## Limitations
|
| 62 |
- Training loss of 1.089 suggests the model is not highly confident in its
|
| 63 |
predictions and will produce suboptimal moves in many positions.
|
| 64 |
+
- QLoRA at 0.49% of parameters means only a small fraction of the model was
|
| 65 |
updated; chess knowledge is partially constrained by the base LLM's pretraining.
|
| 66 |
- Training data consists primarily of high-level Lichess games (depth ≥ 16),
|
| 67 |
meaning the model is tuned on strong engine moves and may struggle with
|