File size: 2,969 Bytes
7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf f2cd858 7dfe6cf | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | ---
tags:
- creativityneuro
- llm-creativity
- mechanistic-interpretability
- arxiv:2607.01433
base_model: microsoft/Phi-3.5-mini-instruct
license: apache-2.0
---
# Phi-3.5-mini · CreativityNeuro
A **CreativityNeuro (CN)** variant of [microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct), with the
weight edit already applied. It loads and runs exactly like the base model.
CreativityNeuro amplifies the parameters that matter for divergent generation but not for
convergent generation, improving divergent thinking with no fine-tuning, no prompt changes,
and no decoding changes.
📄 [Paper](https://arxiv.org/abs/2607.01433) · 💻 [Code](https://github.com/samjschapiro/creativityneuro) · 🤗 [All optimal configs](https://huggingface.co/collections/creativityschapiro/creativityneuro-optimal-configs)
## Configuration
| Parameter | Value |
|---|---|
| Base model | [microsoft/Phi-3.5-mini-instruct](https://huggingface.co/microsoft/Phi-3.5-mini-instruct) |
| ρ (keep ratio) | 0.10 |
| α (amplification) | 1.0 |
| Contrastive prompt set | `dat` |
| Mode | creative |
This is the best-performing CreativityNeuro configuration for Phi-3.5-mini.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("creativityschapiro/phi-3.5-mini-instruct-cn-dat-kr0.1-a1.0-creative")
tokenizer = AutoTokenizer.from_pretrained("creativityschapiro/phi-3.5-mini-instruct-cn-dat-kr0.1-a1.0-creative")
outputs = model.generate(...)
```
## Method
Parameter importance is scored Wanda-style, `S_ij = Σ_b |W_ij| · ‖X_j‖₂`, under two
contrastive prompt sets. The top ρ of each is taken, and the set difference — important for
divergent generation, not for convergent generation — is amplified:
```
W_new = W × (1 + α × mask)
```
To build masks yourself, or apply CN to a model not published here, see
[samjschapiro/creativityneuro](https://github.com/samjschapiro/creativityneuro).
## Results
Across six instruction-tuned models, CreativityNeuro improves scores on the Divergent
Association Task and transfers to open-ended creativity tasks judged by human raters
(N = 720) — the Alternative Uses Test and the Task Task — with gains in originality
(avg. Cohen's *d* = +0.36 AUT, +0.40 TT) and surprise (+0.43 AUT). Full results in the
[paper](https://arxiv.org/abs/2607.01433).
## Citation
```bibtex
@inproceedings{schapiro2026creativityneuro,
title = {CreativityNeuro: Steering Language Model Weights to Improve
Divergent Thinking and Reduce Mode Collapse},
author = {Schapiro, Samuel and Park, Core Francisco and Sosa, Felix
and Varshney, Lav R.},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026},
eprint = {2607.01433},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2607.01433}
}
```
|