Update model card: correct citation, fix load path, link paper and code
Browse files
README.md
CHANGED
|
@@ -3,55 +3,79 @@ tags:
|
|
| 3 |
- creativityneuro
|
| 4 |
- llm-creativity
|
| 5 |
- mechanistic-interpretability
|
|
|
|
| 6 |
base_model: microsoft/Phi-3.5-mini-instruct
|
| 7 |
license: apache-2.0
|
| 8 |
---
|
| 9 |
|
| 10 |
-
#
|
| 11 |
|
| 12 |
-
|
|
|
|
| 13 |
|
| 14 |
-
|
|
|
|
|
|
|
| 15 |
|
| 16 |
-
|
| 17 |
-
- **Modification**: CreativityNeuro weight scaling
|
| 18 |
-
- **Prompt Set**: dat
|
| 19 |
-
- **Keep Ratio**: 0.1 (top 10.0% of task-specific weights)
|
| 20 |
-
- **Alpha**: 1.0 (scaling strength)
|
| 21 |
-
- **Mode**: creative
|
| 22 |
|
| 23 |
-
##
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
-
|
| 29 |
-
W_new = W × (1 + α × mask)
|
| 30 |
-
```
|
| 31 |
-
|
| 32 |
-
Where `mask` identifies weights important for creative tasks but not for routine/associative tasks.
|
| 33 |
|
| 34 |
## Usage
|
| 35 |
|
| 36 |
```python
|
| 37 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 38 |
|
| 39 |
-
model = AutoModelForCausalLM.from_pretrained("
|
| 40 |
-
tokenizer = AutoTokenizer.from_pretrained("
|
| 41 |
|
| 42 |
-
# Use like any other model
|
| 43 |
outputs = model.generate(...)
|
| 44 |
```
|
| 45 |
|
| 46 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
|
| 48 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 49 |
|
| 50 |
```bibtex
|
| 51 |
-
@
|
| 52 |
-
title={CreativityNeuro:
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 56 |
}
|
| 57 |
```
|
|
|
|
| 3 |
- creativityneuro
|
| 4 |
- llm-creativity
|
| 5 |
- mechanistic-interpretability
|
| 6 |
+
- arxiv:2607.01433
|
| 7 |
base_model: microsoft/Phi-3.5-mini-instruct
|
| 8 |
license: apache-2.0
|
| 9 |
---
|
| 10 |
|
| 11 |
+
# Phi-3.5-mini · CreativityNeuro
|
| 12 |
|
| 13 |
+
A **CreativityNeuro (CN)** variant of [microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct), with the
|
| 14 |
+
weight edit already applied. It loads and runs exactly like the base model.
|
| 15 |
|
| 16 |
+
CreativityNeuro amplifies the parameters that matter for divergent generation but not for
|
| 17 |
+
convergent generation, improving divergent thinking with no fine-tuning, no prompt changes,
|
| 18 |
+
and no decoding changes.
|
| 19 |
|
| 20 |
+
📄 [Paper](https://arxiv.org/abs/2607.01433) · 💻 [Code](https://github.com/samjschapiro/creativityneuro) · 🤗 [All optimal configs](https://huggingface.co/collections/creativityschapiro/creativityneuro-optimal-configs)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
|
| 22 |
+
## Configuration
|
| 23 |
|
| 24 |
+
| Parameter | Value |
|
| 25 |
+
|---|---|
|
| 26 |
+
| Base model | [microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct) |
|
| 27 |
+
| ρ (keep ratio) | 0.10 |
|
| 28 |
+
| α (amplification) | 1.0 |
|
| 29 |
+
| Contrastive prompt set | `dat` |
|
| 30 |
+
| Mode | creative |
|
| 31 |
|
| 32 |
+
This is the best-performing CreativityNeuro configuration for Phi-3.5-mini.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 33 |
|
| 34 |
## Usage
|
| 35 |
|
| 36 |
```python
|
| 37 |
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 38 |
|
| 39 |
+
model = AutoModelForCausalLM.from_pretrained("creativityschapiro/phi-3.5-mini-instruct-cn-dat-kr0.1-a1.0-creative")
|
| 40 |
+
tokenizer = AutoTokenizer.from_pretrained("creativityschapiro/phi-3.5-mini-instruct-cn-dat-kr0.1-a1.0-creative")
|
| 41 |
|
|
|
|
| 42 |
outputs = model.generate(...)
|
| 43 |
```
|
| 44 |
|
| 45 |
+
## Method
|
| 46 |
+
|
| 47 |
+
Parameter importance is scored Wanda-style, `S_ij = Σ_b |W_ij| · ‖X_j‖₂`, under two
|
| 48 |
+
contrastive prompt sets. The top ρ of each is taken, and the set difference — important for
|
| 49 |
+
divergent generation, not for convergent generation — is amplified:
|
| 50 |
+
|
| 51 |
+
```
|
| 52 |
+
W_new = W × (1 + α × mask)
|
| 53 |
+
```
|
| 54 |
+
|
| 55 |
+
To build masks yourself, or apply CN to a model not published here, see
|
| 56 |
+
[samjschapiro/creativityneuro](https://github.com/samjschapiro/creativityneuro).
|
| 57 |
|
| 58 |
+
## Results
|
| 59 |
+
|
| 60 |
+
Across six instruction-tuned models, CreativityNeuro improves scores on the Divergent
|
| 61 |
+
Association Task and transfers to open-ended creativity tasks judged by human raters
|
| 62 |
+
(N = 720) — the Alternative Uses Test and the Task Task — with gains in originality
|
| 63 |
+
(avg. Cohen's *d* = +0.36 AUT, +0.40 TT) and surprise (+0.43 AUT). Full results in the
|
| 64 |
+
[paper](https://arxiv.org/abs/2607.01433).
|
| 65 |
+
|
| 66 |
+
## Citation
|
| 67 |
|
| 68 |
```bibtex
|
| 69 |
+
@inproceedings{schapiro2026creativityneuro,
|
| 70 |
+
title = {CreativityNeuro: Steering Language Model Weights to Improve
|
| 71 |
+
Divergent Thinking and Reduce Mode Collapse},
|
| 72 |
+
author = {Schapiro, Samuel and Park, Core Francisco and Sosa, Felix
|
| 73 |
+
and Varshney, Lav R.},
|
| 74 |
+
booktitle = {Conference on Language Modeling (COLM)},
|
| 75 |
+
year = {2026},
|
| 76 |
+
eprint = {2607.01433},
|
| 77 |
+
archivePrefix = {arXiv},
|
| 78 |
+
primaryClass = {cs.AI},
|
| 79 |
+
url = {https://arxiv.org/abs/2607.01433}
|
| 80 |
}
|
| 81 |
```
|