File size: 2,928 Bytes
75ffb5b
 
 
 
 
c4e21cd
75ffb5b
 
 
 
c4e21cd
75ffb5b
c4e21cd
 
75ffb5b
c4e21cd
 
 
75ffb5b
c4e21cd
75ffb5b
c4e21cd
75ffb5b
c4e21cd
 
 
 
 
 
 
75ffb5b
c4e21cd
75ffb5b
 
 
 
 
 
c4e21cd
 
75ffb5b
 
 
 
c4e21cd
 
 
 
 
 
 
 
 
 
 
 
75ffb5b
c4e21cd
 
 
 
 
 
 
 
 
75ffb5b
 
c4e21cd
 
 
 
 
 
 
 
 
 
 
75ffb5b
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
---
tags:
- creativityneuro
- llm-creativity
- mechanistic-interpretability
- arxiv:2607.01433
base_model: Qwen/Qwen2.5-7B-Instruct
license: apache-2.0
---

# Qwen-2.5-7B · CreativityNeuro

A **CreativityNeuro (CN)** variant of [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct), with the
weight edit already applied. It loads and runs exactly like the base model.

CreativityNeuro amplifies the parameters that matter for divergent generation but not for
convergent generation, improving divergent thinking with no fine-tuning, no prompt changes,
and no decoding changes.

📄 [Paper](https://arxiv.org/abs/2607.01433) · 💻 [Code](https://github.com/samjschapiro/creativityneuro) · 🤗 [All optimal configs](https://huggingface.co/collections/creativityschapiro/creativityneuro-optimal-configs)

## Configuration

| Parameter | Value |
|---|---|
| Base model | [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) |
| ρ (keep ratio) | 0.10 |
| α (amplification) | 1.0 |
| Contrastive prompt set | `dat` |
| Mode | creative |

This is the best-performing CreativityNeuro configuration for Qwen-2.5-7B.

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("creativityschapiro/qwen2.5-7b-instruct-cn-dat-kr0.1-a1.0-creative")
tokenizer = AutoTokenizer.from_pretrained("creativityschapiro/qwen2.5-7b-instruct-cn-dat-kr0.1-a1.0-creative")

outputs = model.generate(...)
```

## Method

Parameter importance is scored Wanda-style, `S_ij = Σ_b |W_ij| · ‖X_j‖₂`, under two
contrastive prompt sets. The top ρ of each is taken, and the set difference — important for
divergent generation, not for convergent generation — is amplified:

```
W_new = W × (1 + α × mask)
```

To build masks yourself, or apply CN to a model not published here, see
[samjschapiro/creativityneuro](https://github.com/samjschapiro/creativityneuro).

## Results

Across six instruction-tuned models, CreativityNeuro improves scores on the Divergent
Association Task and transfers to open-ended creativity tasks judged by human raters
(N = 720) — the Alternative Uses Test and the Task Task — with gains in originality
(avg. Cohen's *d* = +0.36 AUT, +0.40 TT) and surprise (+0.43 AUT). Full results in the
[paper](https://arxiv.org/abs/2607.01433).

## Citation

```bibtex
@inproceedings{schapiro2026creativityneuro,
  title         = {CreativityNeuro: Steering Language Model Weights to Improve
                   Divergent Thinking and Reduce Mode Collapse},
  author        = {Schapiro, Samuel and Park, Core Francisco and Sosa, Felix
                   and Varshney, Lav R.},
  booktitle     = {Conference on Language Modeling (COLM)},
  year          = {2026},
  eprint        = {2607.01433},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2607.01433}
}
```