Text Generation
Transformers
recurrent_qwen
recurrent-depth
latent-reasoning
qwen2.5
research
custom_code
Instructions to use mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper
- SGLang
How to use mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper with Docker Model Runner:
docker model run hf.co/mshapiro123/recurrent-qwen2.5-0.5b-natural-keeper
Prepare private Paper One model release
Browse files- .gitattributes +4 -34
- LICENSE +202 -0
- README.md +41 -0
- config.json +29 -0
- configuration_recurrent_qwen.py +47 -0
- conversion_receipt.json +43 -0
- modeling_recurrent_qwen.py +486 -0
- recurrent_delta.safetensors +3 -0
- verification_spec.json +14 -0
- verification_subset.jsonl +8 -0
.gitattributes
CHANGED
|
@@ -1,35 +1,5 @@
|
|
| 1 |
-
*.7z filter=lfs diff=lfs merge=lfs -text
|
| 2 |
-
*.arrow filter=lfs diff=lfs merge=lfs -text
|
| 3 |
-
*.bin filter=lfs diff=lfs merge=lfs -text
|
| 4 |
-
*.bz2 filter=lfs diff=lfs merge=lfs -text
|
| 5 |
-
*.ckpt filter=lfs diff=lfs merge=lfs -text
|
| 6 |
-
*.ftz filter=lfs diff=lfs merge=lfs -text
|
| 7 |
-
*.gz filter=lfs diff=lfs merge=lfs -text
|
| 8 |
-
*.h5 filter=lfs diff=lfs merge=lfs -text
|
| 9 |
-
*.joblib filter=lfs diff=lfs merge=lfs -text
|
| 10 |
-
*.lfs.* filter=lfs diff=lfs merge=lfs -text
|
| 11 |
-
*.mlmodel filter=lfs diff=lfs merge=lfs -text
|
| 12 |
-
*.model filter=lfs diff=lfs merge=lfs -text
|
| 13 |
-
*.msgpack filter=lfs diff=lfs merge=lfs -text
|
| 14 |
-
*.npy filter=lfs diff=lfs merge=lfs -text
|
| 15 |
-
*.npz filter=lfs diff=lfs merge=lfs -text
|
| 16 |
-
*.onnx filter=lfs diff=lfs merge=lfs -text
|
| 17 |
-
*.ot filter=lfs diff=lfs merge=lfs -text
|
| 18 |
-
*.parquet filter=lfs diff=lfs merge=lfs -text
|
| 19 |
-
*.pb filter=lfs diff=lfs merge=lfs -text
|
| 20 |
-
*.pickle filter=lfs diff=lfs merge=lfs -text
|
| 21 |
-
*.pkl filter=lfs diff=lfs merge=lfs -text
|
| 22 |
-
*.pt filter=lfs diff=lfs merge=lfs -text
|
| 23 |
-
*.pth filter=lfs diff=lfs merge=lfs -text
|
| 24 |
-
*.rar filter=lfs diff=lfs merge=lfs -text
|
| 25 |
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 26 |
-
|
| 27 |
-
*.
|
| 28 |
-
*.
|
| 29 |
-
|
| 30 |
-
*.tgz filter=lfs diff=lfs merge=lfs -text
|
| 31 |
-
*.wasm filter=lfs diff=lfs merge=lfs -text
|
| 32 |
-
*.xz filter=lfs diff=lfs merge=lfs -text
|
| 33 |
-
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
-
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
-
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
*.safetensors filter=lfs diff=lfs merge=lfs -text
|
| 2 |
+
*.json text eol=lf
|
| 3 |
+
*.md text eol=lf
|
| 4 |
+
*.py text eol=lf
|
| 5 |
+
LICENSE text eol=lf
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
LICENSE
ADDED
|
@@ -0,0 +1,202 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Apache License
|
| 2 |
+
Version 2.0, January 2004
|
| 3 |
+
http://www.apache.org/licenses/
|
| 4 |
+
|
| 5 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 6 |
+
|
| 7 |
+
1. Definitions.
|
| 8 |
+
|
| 9 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 10 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 11 |
+
|
| 12 |
+
"Licensor" shall mean the copyright owner or entity authorized by
|
| 13 |
+
the copyright owner that is granting the License.
|
| 14 |
+
|
| 15 |
+
"Legal Entity" shall mean the union of the acting entity and all
|
| 16 |
+
other entities that control, are controlled by, or are under common
|
| 17 |
+
control with that entity. For the purposes of this definition,
|
| 18 |
+
"control" means (i) the power, direct or indirect, to cause the
|
| 19 |
+
direction or management of such entity, whether by contract or
|
| 20 |
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
| 21 |
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
| 22 |
+
|
| 23 |
+
"You" (or "Your") shall mean an individual or Legal Entity
|
| 24 |
+
exercising permissions granted by this License.
|
| 25 |
+
|
| 26 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 27 |
+
including but not limited to software source code, documentation
|
| 28 |
+
source, and configuration files.
|
| 29 |
+
|
| 30 |
+
"Object" form shall mean any form resulting from mechanical
|
| 31 |
+
transformation or translation of a Source form, including but
|
| 32 |
+
not limited to compiled object code, generated documentation,
|
| 33 |
+
and conversions to other media types.
|
| 34 |
+
|
| 35 |
+
"Work" shall mean the work of authorship, whether in Source or
|
| 36 |
+
Object form, made available under the License, as indicated by a
|
| 37 |
+
copyright notice that is included in or attached to the work
|
| 38 |
+
(an example is provided in the Appendix below).
|
| 39 |
+
|
| 40 |
+
"Derivative Works" shall mean any work, whether in Source or Object
|
| 41 |
+
form, that is based on (or derived from) the Work and for which the
|
| 42 |
+
editorial revisions, annotations, elaborations, or other modifications
|
| 43 |
+
represent, as a whole, an original work of authorship. For the purposes
|
| 44 |
+
of this License, Derivative Works shall not include works that remain
|
| 45 |
+
separable from, or merely link (or bind by name) to the interfaces of,
|
| 46 |
+
the Work and Derivative Works thereof.
|
| 47 |
+
|
| 48 |
+
"Contribution" shall mean any work of authorship, including
|
| 49 |
+
the original version of the Work and any modifications or additions
|
| 50 |
+
to that Work or Derivative Works thereof, that is intentionally
|
| 51 |
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
| 52 |
+
or by an individual or Legal Entity authorized to submit on behalf of
|
| 53 |
+
the copyright owner. For the purposes of this definition, "submitted"
|
| 54 |
+
means any form of electronic, verbal, or written communication sent
|
| 55 |
+
to the Licensor or its representatives, including but not limited to
|
| 56 |
+
communication on electronic mailing lists, source code control systems,
|
| 57 |
+
and issue tracking systems that are managed by, or on behalf of, the
|
| 58 |
+
Licensor for the purpose of discussing and improving the Work, but
|
| 59 |
+
excluding communication that is conspicuously marked or otherwise
|
| 60 |
+
designated in writing by the copyright owner as "Not a Contribution."
|
| 61 |
+
|
| 62 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
| 63 |
+
on behalf of whom a Contribution has been received by Licensor and
|
| 64 |
+
subsequently incorporated within the Work.
|
| 65 |
+
|
| 66 |
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
| 67 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 68 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 69 |
+
copyright license to reproduce, prepare Derivative Works of,
|
| 70 |
+
publicly display, publicly perform, sublicense, and distribute the
|
| 71 |
+
Work and such Derivative Works in Source or Object form.
|
| 72 |
+
|
| 73 |
+
3. Grant of Patent License. Subject to the terms and conditions of
|
| 74 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 75 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 76 |
+
(except as stated in this section) patent license to make, have made,
|
| 77 |
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
| 78 |
+
where such license applies only to those patent claims licensable
|
| 79 |
+
by such Contributor that are necessarily infringed by their
|
| 80 |
+
Contribution(s) alone or by combination of their Contribution(s)
|
| 81 |
+
with the Work to which such Contribution(s) was submitted. If You
|
| 82 |
+
institute patent litigation against any entity (including a
|
| 83 |
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
| 84 |
+
or a Contribution incorporated within the Work constitutes direct
|
| 85 |
+
or contributory patent infringement, then any patent licenses
|
| 86 |
+
granted to You under this License for that Work shall terminate
|
| 87 |
+
as of the date such litigation is filed.
|
| 88 |
+
|
| 89 |
+
4. Redistribution. You may reproduce and distribute copies of the
|
| 90 |
+
Work or Derivative Works thereof in any medium, with or without
|
| 91 |
+
modifications, and in Source or Object form, provided that You
|
| 92 |
+
meet the following conditions:
|
| 93 |
+
|
| 94 |
+
(a) You must give any other recipients of the Work or
|
| 95 |
+
Derivative Works a copy of this License; and
|
| 96 |
+
|
| 97 |
+
(b) You must cause any modified files to carry prominent notices
|
| 98 |
+
stating that You changed the files; and
|
| 99 |
+
|
| 100 |
+
(c) You must retain, in the Source form of any Derivative Works
|
| 101 |
+
that You distribute, all copyright, patent, trademark, and
|
| 102 |
+
attribution notices from the Source form of the Work,
|
| 103 |
+
excluding those notices that do not pertain to any part of
|
| 104 |
+
the Derivative Works; and
|
| 105 |
+
|
| 106 |
+
(d) If the Work includes a "NOTICE" text file as part of its
|
| 107 |
+
distribution, then any Derivative Works that You distribute must
|
| 108 |
+
include a readable copy of the attribution notices contained
|
| 109 |
+
within such NOTICE file, excluding those notices that do not
|
| 110 |
+
pertain to any part of the Derivative Works, in at least one
|
| 111 |
+
of the following places: within a NOTICE text file distributed
|
| 112 |
+
as part of the Derivative Works; within the Source form or
|
| 113 |
+
documentation, if provided along with the Derivative Works; or,
|
| 114 |
+
within a display generated by the Derivative Works, if and
|
| 115 |
+
wherever such third-party notices normally appear. The contents
|
| 116 |
+
of the NOTICE file are for informational purposes only and
|
| 117 |
+
do not modify the License. You may add Your own attribution
|
| 118 |
+
notices within Derivative Works that You distribute, alongside
|
| 119 |
+
or as an addendum to the NOTICE text from the Work, provided
|
| 120 |
+
that such additional attribution notices cannot be construed
|
| 121 |
+
as modifying the License.
|
| 122 |
+
|
| 123 |
+
You may add Your own copyright statement to Your modifications and
|
| 124 |
+
may provide additional or different license terms and conditions
|
| 125 |
+
for use, reproduction, or distribution of Your modifications, or
|
| 126 |
+
for any such Derivative Works as a whole, provided Your use,
|
| 127 |
+
reproduction, and distribution of the Work otherwise complies with
|
| 128 |
+
the conditions stated in this License.
|
| 129 |
+
|
| 130 |
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
| 131 |
+
any Contribution intentionally submitted for inclusion in the Work
|
| 132 |
+
by You to the Licensor shall be under the terms and conditions of
|
| 133 |
+
this License, without any additional terms or conditions.
|
| 134 |
+
Notwithstanding the above, nothing herein shall supersede or modify
|
| 135 |
+
the terms of any separate license agreement you may have executed
|
| 136 |
+
with Licensor regarding such Contributions.
|
| 137 |
+
|
| 138 |
+
6. Trademarks. This License does not grant permission to use the trade
|
| 139 |
+
names, trademarks, service marks, or product names of the Licensor,
|
| 140 |
+
except as required for reasonable and customary use in describing the
|
| 141 |
+
origin of the Work and reproducing the content of the NOTICE file.
|
| 142 |
+
|
| 143 |
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
| 144 |
+
agreed to in writing, Licensor provides the Work (and each
|
| 145 |
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
| 146 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
| 147 |
+
implied, including, without limitation, any warranties or conditions
|
| 148 |
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 149 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 150 |
+
appropriateness of using or redistributing the Work and assume any
|
| 151 |
+
risks associated with Your exercise of permissions under this License.
|
| 152 |
+
|
| 153 |
+
8. Limitation of Liability. In no event and under no legal theory,
|
| 154 |
+
whether in tort (including negligence), contract, or otherwise,
|
| 155 |
+
unless required by applicable law (such as deliberate and grossly
|
| 156 |
+
negligent acts) or agreed to in writing, shall any Contributor be
|
| 157 |
+
liable to You for damages, including any direct, indirect, special,
|
| 158 |
+
incidental, or consequential damages of any character arising as a
|
| 159 |
+
result of this License or out of the use or inability to use the
|
| 160 |
+
Work (including but not limited to damages for loss of goodwill,
|
| 161 |
+
work stoppage, computer failure or malfunction, or any and all
|
| 162 |
+
other commercial damages or losses), even if such Contributor
|
| 163 |
+
has been advised of the possibility of such damages.
|
| 164 |
+
|
| 165 |
+
9. Accepting Warranty or Additional Liability. While redistributing
|
| 166 |
+
the Work or Derivative Works thereof, You may choose to offer,
|
| 167 |
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
| 168 |
+
or other liability obligations and/or rights consistent with this
|
| 169 |
+
License. However, in accepting such obligations, You may act only
|
| 170 |
+
on Your own behalf and on Your sole responsibility, not on behalf
|
| 171 |
+
of any other Contributor, and only if You agree to indemnify,
|
| 172 |
+
defend, and hold each Contributor harmless for any liability
|
| 173 |
+
incurred by, or claims asserted against, such Contributor by reason
|
| 174 |
+
of your accepting any such warranty or additional liability.
|
| 175 |
+
|
| 176 |
+
END OF TERMS AND CONDITIONS
|
| 177 |
+
|
| 178 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 179 |
+
|
| 180 |
+
To apply the Apache License to your work, attach the following
|
| 181 |
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
| 182 |
+
replaced with your own identifying information. (Don't include
|
| 183 |
+
the brackets!) The text should be enclosed in the appropriate
|
| 184 |
+
comment syntax for the file format. We also recommend that a
|
| 185 |
+
file or class name and description of purpose be included on the
|
| 186 |
+
same "printed page" as the copyright notice for easier
|
| 187 |
+
identification within third-party archives.
|
| 188 |
+
|
| 189 |
+
Copyright 2026 Mark Shapiro
|
| 190 |
+
|
| 191 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 192 |
+
you may not use this file except in compliance with the License.
|
| 193 |
+
You may obtain a copy of the License at
|
| 194 |
+
|
| 195 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 196 |
+
|
| 197 |
+
Unless required by applicable law or agreed to in writing, software
|
| 198 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 199 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 200 |
+
See the License for the specific language governing permissions and
|
| 201 |
+
limitations under the License.
|
| 202 |
+
|
README.md
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: Qwen/Qwen2.5-0.5B-Instruct
|
| 4 |
+
library_name: transformers
|
| 5 |
+
tags:
|
| 6 |
+
- recurrent-depth
|
| 7 |
+
- latent-reasoning
|
| 8 |
+
- qwen2.5
|
| 9 |
+
- research
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# recurrent-qwen2.5-0.5b-natural-keeper
|
| 13 |
+
|
| 14 |
+
The frozen natural keeper of the full-block recurrent retrofit: the verbally trained step-2,000 checkpoint from *Retrofitting Recurrent Depth into a Pretrained Language Model* (arXiv:2608.11233), and the substrate designated for the registered companion study on guided stochastic width.
|
| 15 |
+
|
| 16 |
+
## What this checkpoint is
|
| 17 |
+
|
| 18 |
+
The same full-block architecture as the companion full-block release (Prelude 0-5, weight-tied Recurrent Block 6-17 executed T times, Coda 18-23, split re-entry bridge; 180,556,929 forward-active trained parameters; T = 1 reproduces the base computation exactly; loads with `trust_remote_code=True`, using the modeling code shipped in this repository). After verbal fine-tuning on generated relay surfaces, checkpoints past step 2,000 kept improving aggregate accuracy while the worst-case deep tail contracted monotonically. This checkpoint was frozen at the deep-tail peak: the worst-case deep-tail minimum across the two verbal surfaces peaked here at 54.69%, against 19.53% by step 6,000. It was selected to preserve worst-case deep behavior rather than to win on aggregate accuracy.
|
| 19 |
+
|
| 20 |
+
## Usage contract
|
| 21 |
+
|
| 22 |
+
Forced-depth evaluation only: loops = task depth, no learned halting. General use should run at T = 1. The keeper lineage is preregistered non-inferior to the base model on ARC-Easy and ARC-Challenge at loop 1.
|
| 23 |
+
|
| 24 |
+
## Measured results (receipts in the companion repository)
|
| 25 |
+
|
| 26 |
+
At step 6,000 the same lineage reached 86.0% (relay) and 79.0% (pointer) on the controlled verbal surfaces; this step-2,000 keeper trades aggregate accuracy for the strongest deep tail. It passed the branching-relations validity screen, 389/512 (75.98%) pooled with a minimum depth accuracy of 62.5%, which qualifies it as the substrate for the companion stochastic-width study. Verbal competence comes from verbal training: without it, zero-shot transfer to these surfaces was 16-20% at both budgets.
|
| 27 |
+
|
| 28 |
+
## Limitations — claims this model does not support
|
| 29 |
+
|
| 30 |
+
- The verbal surfaces are templated and distractor-free renderings of the synthetic task, not broad natural reasoning benchmarks. No broad natural-language reasoning gains are claimed.
|
| 31 |
+
- No learned depth selection and no stochastic-width claims. The width screen establishes substrate competence only; no latent prior or posterior head was trained.
|
| 32 |
+
- Preservation is battery-scoped non-inferiority at loop 1, not universal.
|
| 33 |
+
- Continued training toward new operations breaches the acquisition-retention boundary documented in the paper.
|
| 34 |
+
|
| 35 |
+
## Provenance
|
| 36 |
+
|
| 37 |
+
Checkpoint SHA-256: `0f657b653078ba403cbc666410e7598ca20c836d5bd6e19a0e85a186a82c5d2f`. Receipts, claim ledger, preregistrations: https://github.com/mshapiro123/recurrent-qwen-svgd. Base model: Qwen2.5-0.5B-Instruct (Apache 2.0), Qwen2.5 Technical Report (arXiv:2412.15115).
|
| 38 |
+
|
| 39 |
+
## Citation
|
| 40 |
+
|
| 41 |
+
Shapiro, M. (2026). Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets. arXiv:2608.11233.
|
config.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"RecurrentQwenForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"auto_map": {
|
| 6 |
+
"AutoConfig": "configuration_recurrent_qwen.RecurrentQwenConfig",
|
| 7 |
+
"AutoModelForCausalLM": "modeling_recurrent_qwen.RecurrentQwenForCausalLM"
|
| 8 |
+
},
|
| 9 |
+
"base_model_name_or_path": "Qwen/Qwen2.5-0.5B-Instruct",
|
| 10 |
+
"base_model_revision": "main",
|
| 11 |
+
"checkpoint_kind": "full_block_delta",
|
| 12 |
+
"delta_filename": "recurrent_delta.safetensors",
|
| 13 |
+
"lora_alpha": 32,
|
| 14 |
+
"lora_rank": 16,
|
| 15 |
+
"lora_target_modules": [
|
| 16 |
+
"q_proj",
|
| 17 |
+
"k_proj",
|
| 18 |
+
"v_proj",
|
| 19 |
+
"o_proj",
|
| 20 |
+
"gate_proj",
|
| 21 |
+
"up_proj",
|
| 22 |
+
"down_proj"
|
| 23 |
+
],
|
| 24 |
+
"model_type": "recurrent_qwen",
|
| 25 |
+
"prelude_end": 6,
|
| 26 |
+
"recurrent_end": 18,
|
| 27 |
+
"transformers_version": "4.53.0",
|
| 28 |
+
"use_cache": false
|
| 29 |
+
}
|
configuration_recurrent_qwen.py
ADDED
|
@@ -0,0 +1,47 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Hugging Face configuration for the forced-depth recurrent Qwen release."""
|
| 2 |
+
|
| 3 |
+
from __future__ import annotations
|
| 4 |
+
|
| 5 |
+
from typing import Any
|
| 6 |
+
|
| 7 |
+
from transformers import PreTrainedConfig
|
| 8 |
+
|
| 9 |
+
|
| 10 |
+
class RecurrentQwenConfig(PreTrainedConfig):
|
| 11 |
+
"""Configuration for a Qwen base plus a recurrent-depth delta."""
|
| 12 |
+
|
| 13 |
+
model_type = "recurrent_qwen"
|
| 14 |
+
|
| 15 |
+
def __init__(
|
| 16 |
+
self,
|
| 17 |
+
*,
|
| 18 |
+
base_model_name_or_path: str = "Qwen/Qwen2.5-0.5B-Instruct",
|
| 19 |
+
base_model_revision: str = "main",
|
| 20 |
+
prelude_end: int = 6,
|
| 21 |
+
recurrent_end: int = 18,
|
| 22 |
+
checkpoint_kind: str = "full_block_delta",
|
| 23 |
+
delta_filename: str = "recurrent_delta.safetensors",
|
| 24 |
+
lora_rank: int = 16,
|
| 25 |
+
lora_alpha: int = 32,
|
| 26 |
+
lora_target_modules: list[str] | None = None,
|
| 27 |
+
**kwargs: Any,
|
| 28 |
+
) -> None:
|
| 29 |
+
super().__init__(**kwargs)
|
| 30 |
+
if not 0 < int(prelude_end) < int(recurrent_end):
|
| 31 |
+
raise ValueError("Expected 0 < prelude_end < recurrent_end")
|
| 32 |
+
if checkpoint_kind not in {"full_block_delta", "lora_adapter"}:
|
| 33 |
+
raise ValueError("checkpoint_kind must be full_block_delta or lora_adapter")
|
| 34 |
+
self.base_model_name_or_path = str(base_model_name_or_path)
|
| 35 |
+
self.base_model_revision = str(base_model_revision)
|
| 36 |
+
self.prelude_end = int(prelude_end)
|
| 37 |
+
self.recurrent_end = int(recurrent_end)
|
| 38 |
+
self.checkpoint_kind = str(checkpoint_kind)
|
| 39 |
+
self.delta_filename = str(delta_filename)
|
| 40 |
+
self.lora_rank = int(lora_rank)
|
| 41 |
+
self.lora_alpha = int(lora_alpha)
|
| 42 |
+
self.lora_target_modules = list(
|
| 43 |
+
lora_target_modules
|
| 44 |
+
or ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"]
|
| 45 |
+
)
|
| 46 |
+
self.use_cache = False
|
| 47 |
+
|
conversion_receipt.json
ADDED
|
@@ -0,0 +1,43 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"schema_version": 1,
|
| 3 |
+
"kind": "paper_one_hf_checkpoint_conversion",
|
| 4 |
+
"status": "green",
|
| 5 |
+
"created_at_utc": "2026-08-16T04:39:04.122744Z",
|
| 6 |
+
"repo_name": "recurrent-qwen2.5-0.5b-natural-keeper",
|
| 7 |
+
"checkpoint_kind": "full_block_delta",
|
| 8 |
+
"source_checkpoint": "/content/drive/MyDrive/recurrent-qwen-svgd-backups/natural_surface_backup_20260709_180835/checkpoints/stage5_natural_surface_transfer_rung0_fixed_prompt_20260709_133812/unfrozen_recurrent_step_2000.pt",
|
| 9 |
+
"source_checkpoint_sha256_expected": "0f657b653078ba403cbc666410e7598ca20c836d5bd6e19a0e85a186a82c5d2f",
|
| 10 |
+
"source_checkpoint_sha256": "0f657b653078ba403cbc666410e7598ca20c836d5bd6e19a0e85a186a82c5d2f",
|
| 11 |
+
"source_phase": "unfrozen_recurrent",
|
| 12 |
+
"source_step": 2000,
|
| 13 |
+
"source_tensor_count": 152,
|
| 14 |
+
"source_total_parameters": 182163457,
|
| 15 |
+
"excluded_checkpoint_tensors": {
|
| 16 |
+
"bridge.proj.bias": {
|
| 17 |
+
"shape": [
|
| 18 |
+
896
|
| 19 |
+
],
|
| 20 |
+
"parameters": 896
|
| 21 |
+
},
|
| 22 |
+
"bridge.proj.weight": {
|
| 23 |
+
"shape": [
|
| 24 |
+
896,
|
| 25 |
+
1792
|
| 26 |
+
],
|
| 27 |
+
"parameters": 1605632
|
| 28 |
+
}
|
| 29 |
+
},
|
| 30 |
+
"excluded_checkpoint_parameters": 1606528,
|
| 31 |
+
"exclusion_reason": "Receipt-bound legacy concat projection bypassed by split-mode forward execution",
|
| 32 |
+
"safetensors_file": "recurrent_delta.safetensors",
|
| 33 |
+
"safetensors_sha256": "a15147dc4338c2ba21b3dc4ec440824516344153fe16bd95ac33e3d726c6d1c0",
|
| 34 |
+
"tensor_count": 150,
|
| 35 |
+
"total_parameters": 180556929,
|
| 36 |
+
"lora_parameters": 0,
|
| 37 |
+
"bridge_parameters": 1608321,
|
| 38 |
+
"dtype_tensor_counts": {
|
| 39 |
+
"bfloat16": 144,
|
| 40 |
+
"float32": 6
|
| 41 |
+
},
|
| 42 |
+
"key_transform": "exclude manifest-listed inactive compatibility tensors; base_model.* -> backbone.*; all other release keys unchanged"
|
| 43 |
+
}
|
modeling_recurrent_qwen.py
ADDED
|
@@ -0,0 +1,486 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Self-contained forced-depth recurrent Qwen loader for the Paper One release.
|
| 2 |
+
|
| 3 |
+
This module deliberately contains no halting head or adaptive-depth path. Every
|
| 4 |
+
forward call uses an externally supplied ``max_loops`` and, optionally, an
|
| 5 |
+
externally supplied per-row ``loop_selection`` bounded by that maximum.
|
| 6 |
+
"""
|
| 7 |
+
|
| 8 |
+
from __future__ import annotations
|
| 9 |
+
|
| 10 |
+
import inspect
|
| 11 |
+
from dataclasses import dataclass
|
| 12 |
+
from pathlib import Path
|
| 13 |
+
from typing import Any, Optional
|
| 14 |
+
|
| 15 |
+
import torch
|
| 16 |
+
import torch.nn.functional as F
|
| 17 |
+
from safetensors.torch import load_file
|
| 18 |
+
from torch import nn
|
| 19 |
+
from transformers import AutoModelForCausalLM, PreTrainedModel
|
| 20 |
+
from transformers.generation import GenerationMixin
|
| 21 |
+
from transformers.modeling_outputs import CausalLMOutputWithPast
|
| 22 |
+
from transformers.utils.hub import cached_file
|
| 23 |
+
|
| 24 |
+
from .configuration_recurrent_qwen import RecurrentQwenConfig
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
@dataclass
|
| 28 |
+
class RecurrentCausalLMOutput(CausalLMOutputWithPast):
|
| 29 |
+
"""Causal-LM output with optional loop-indexed logits."""
|
| 30 |
+
|
| 31 |
+
loop_logits: Optional[torch.FloatTensor] = None
|
| 32 |
+
selected_loop_counts: Optional[torch.LongTensor] = None
|
| 33 |
+
|
| 34 |
+
|
| 35 |
+
class IdentityGatedBridge(nn.Module):
|
| 36 |
+
"""Split re-entry bridge used by the frozen Paper One checkpoints."""
|
| 37 |
+
|
| 38 |
+
def __init__(self, hidden_size: int) -> None:
|
| 39 |
+
super().__init__()
|
| 40 |
+
self.hidden_size = int(hidden_size)
|
| 41 |
+
self.prelude_norm = nn.LayerNorm(hidden_size)
|
| 42 |
+
self.prelude_proj = nn.Linear(hidden_size, hidden_size, bias=False)
|
| 43 |
+
self.state_proj = nn.Linear(hidden_size, hidden_size, bias=True)
|
| 44 |
+
self.bridge_gate = nn.Parameter(torch.tensor(1.0))
|
| 45 |
+
with torch.no_grad():
|
| 46 |
+
self.prelude_proj.weight.zero_()
|
| 47 |
+
self.state_proj.weight.copy_(torch.eye(hidden_size))
|
| 48 |
+
self.state_proj.bias.zero_()
|
| 49 |
+
|
| 50 |
+
def forward(self, state: torch.Tensor, prelude: torch.Tensor) -> torch.Tensor:
|
| 51 |
+
input_dtype = state.dtype
|
| 52 |
+
work_dtype = self.state_proj.weight.dtype
|
| 53 |
+
work = state.to(dtype=work_dtype)
|
| 54 |
+
normalized_prelude = self.prelude_norm(prelude.to(dtype=work_dtype))
|
| 55 |
+
translated = self.prelude_proj(normalized_prelude) + self.state_proj(work)
|
| 56 |
+
gate = self.bridge_gate.to(device=state.device, dtype=work_dtype)
|
| 57 |
+
return (work + gate * (translated - work)).to(dtype=input_dtype)
|
| 58 |
+
|
| 59 |
+
|
| 60 |
+
class LoRALinear(nn.Module):
|
| 61 |
+
"""Inference-only rank-decomposition wrapper matching the training keys."""
|
| 62 |
+
|
| 63 |
+
def __init__(self, base: nn.Linear, *, rank: int, alpha: int) -> None:
|
| 64 |
+
super().__init__()
|
| 65 |
+
self.base = base
|
| 66 |
+
self.rank = int(rank)
|
| 67 |
+
self.alpha = int(alpha)
|
| 68 |
+
self.scaling = float(alpha) / float(rank)
|
| 69 |
+
self.lora_a = nn.Linear(base.in_features, rank, bias=False, dtype=torch.float32)
|
| 70 |
+
self.lora_b = nn.Linear(rank, base.out_features, bias=False, dtype=torch.float32)
|
| 71 |
+
with torch.no_grad():
|
| 72 |
+
nn.init.kaiming_uniform_(self.lora_a.weight, a=5**0.5)
|
| 73 |
+
self.lora_b.weight.zero_()
|
| 74 |
+
for parameter in self.base.parameters():
|
| 75 |
+
parameter.requires_grad_(False)
|
| 76 |
+
|
| 77 |
+
def forward(self, inputs: torch.Tensor) -> torch.Tensor:
|
| 78 |
+
base_output = self.base(inputs)
|
| 79 |
+
adapter = self.lora_b(self.lora_a(inputs.float())) * self.scaling
|
| 80 |
+
return base_output + adapter.to(dtype=base_output.dtype)
|
| 81 |
+
|
| 82 |
+
|
| 83 |
+
def _replace_lora_targets(
|
| 84 |
+
module: nn.Module,
|
| 85 |
+
target_names: set[str],
|
| 86 |
+
*,
|
| 87 |
+
rank: int,
|
| 88 |
+
alpha: int,
|
| 89 |
+
) -> int:
|
| 90 |
+
replaced = 0
|
| 91 |
+
for child_name, child in list(module.named_children()):
|
| 92 |
+
if isinstance(child, LoRALinear):
|
| 93 |
+
continue
|
| 94 |
+
if child_name in target_names and isinstance(child, nn.Linear):
|
| 95 |
+
setattr(module, child_name, LoRALinear(child, rank=rank, alpha=alpha))
|
| 96 |
+
replaced += 1
|
| 97 |
+
else:
|
| 98 |
+
replaced += _replace_lora_targets(child, target_names, rank=rank, alpha=alpha)
|
| 99 |
+
return replaced
|
| 100 |
+
|
| 101 |
+
|
| 102 |
+
class RecurrentQwenForCausalLM(PreTrainedModel, GenerationMixin):
|
| 103 |
+
"""Qwen causal LM with a forced, weight-tied middle-block recurrence."""
|
| 104 |
+
|
| 105 |
+
config_class = RecurrentQwenConfig
|
| 106 |
+
base_model_prefix = "backbone"
|
| 107 |
+
main_input_name = "input_ids"
|
| 108 |
+
|
| 109 |
+
def __init__(self, config: RecurrentQwenConfig, backbone: nn.Module) -> None:
|
| 110 |
+
super().__init__(config)
|
| 111 |
+
self.backbone = backbone
|
| 112 |
+
if not hasattr(backbone, "model") or not hasattr(backbone.model, "layers"):
|
| 113 |
+
raise TypeError("Expected a Qwen-style causal LM with .model.layers")
|
| 114 |
+
if int(config.recurrent_end) >= len(self.qwen.layers):
|
| 115 |
+
raise ValueError("recurrent_end must leave at least one coda layer")
|
| 116 |
+
hidden_size = int(getattr(backbone.config, "hidden_size"))
|
| 117 |
+
base_parameter = next(backbone.parameters())
|
| 118 |
+
self.bridge = IdentityGatedBridge(hidden_size).to(
|
| 119 |
+
device=base_parameter.device,
|
| 120 |
+
dtype=base_parameter.dtype,
|
| 121 |
+
)
|
| 122 |
+
self.lora_module_count = 0
|
| 123 |
+
if config.checkpoint_kind == "lora_adapter":
|
| 124 |
+
targets = set(config.lora_target_modules)
|
| 125 |
+
for layer_index in range(config.prelude_end, config.recurrent_end):
|
| 126 |
+
self.lora_module_count += _replace_lora_targets(
|
| 127 |
+
self.qwen.layers[layer_index],
|
| 128 |
+
targets,
|
| 129 |
+
rank=config.lora_rank,
|
| 130 |
+
alpha=config.lora_alpha,
|
| 131 |
+
)
|
| 132 |
+
if self.lora_module_count != 84:
|
| 133 |
+
raise RuntimeError(f"Expected 84 recurrent LoRA modules, got {self.lora_module_count}")
|
| 134 |
+
|
| 135 |
+
@property
|
| 136 |
+
def qwen(self) -> nn.Module:
|
| 137 |
+
return self.backbone.model
|
| 138 |
+
|
| 139 |
+
@property
|
| 140 |
+
def lm_head(self) -> nn.Module:
|
| 141 |
+
return self.backbone.lm_head
|
| 142 |
+
|
| 143 |
+
def get_input_embeddings(self) -> nn.Module:
|
| 144 |
+
return self.qwen.embed_tokens
|
| 145 |
+
|
| 146 |
+
def set_input_embeddings(self, value: nn.Module) -> None:
|
| 147 |
+
self.qwen.embed_tokens = value
|
| 148 |
+
|
| 149 |
+
def get_output_embeddings(self) -> nn.Module:
|
| 150 |
+
return self.lm_head
|
| 151 |
+
|
| 152 |
+
def set_output_embeddings(self, value: nn.Module) -> None:
|
| 153 |
+
self.backbone.lm_head = value
|
| 154 |
+
|
| 155 |
+
@classmethod
|
| 156 |
+
def from_pretrained(
|
| 157 |
+
cls,
|
| 158 |
+
pretrained_model_name_or_path: str | Path,
|
| 159 |
+
*model_args: Any,
|
| 160 |
+
config: RecurrentQwenConfig | None = None,
|
| 161 |
+
**kwargs: Any,
|
| 162 |
+
) -> "RecurrentQwenForCausalLM":
|
| 163 |
+
"""Load the pinned base model, construct the surgery, then apply the delta."""
|
| 164 |
+
|
| 165 |
+
if model_args:
|
| 166 |
+
raise TypeError("Positional model arguments are not supported by this release loader")
|
| 167 |
+
token = kwargs.pop("token", None)
|
| 168 |
+
revision = kwargs.pop("revision", "main")
|
| 169 |
+
cache_dir = kwargs.pop("cache_dir", None)
|
| 170 |
+
local_files_only = bool(kwargs.pop("local_files_only", False))
|
| 171 |
+
kwargs.pop("trust_remote_code", None)
|
| 172 |
+
kwargs.pop("_from_auto", None)
|
| 173 |
+
kwargs.pop("_fast_init", None)
|
| 174 |
+
kwargs.pop("state_dict", None)
|
| 175 |
+
kwargs.pop("weights_only", None)
|
| 176 |
+
kwargs.pop("adapter_kwargs", None)
|
| 177 |
+
|
| 178 |
+
if config is None:
|
| 179 |
+
config = RecurrentQwenConfig.from_pretrained(
|
| 180 |
+
pretrained_model_name_or_path,
|
| 181 |
+
revision=revision,
|
| 182 |
+
token=token,
|
| 183 |
+
cache_dir=cache_dir,
|
| 184 |
+
local_files_only=local_files_only,
|
| 185 |
+
)
|
| 186 |
+
|
| 187 |
+
base_keys = {
|
| 188 |
+
"attn_implementation",
|
| 189 |
+
"device_map",
|
| 190 |
+
"dtype",
|
| 191 |
+
"torch_dtype",
|
| 192 |
+
"low_cpu_mem_usage",
|
| 193 |
+
"max_memory",
|
| 194 |
+
"offload_folder",
|
| 195 |
+
"offload_state_dict",
|
| 196 |
+
}
|
| 197 |
+
base_kwargs = {key: kwargs.pop(key) for key in list(kwargs) if key in base_keys}
|
| 198 |
+
if kwargs:
|
| 199 |
+
unknown = ", ".join(sorted(kwargs))
|
| 200 |
+
raise TypeError(f"Unsupported loader keyword(s): {unknown}")
|
| 201 |
+
base_kwargs.update(
|
| 202 |
+
{
|
| 203 |
+
"revision": config.base_model_revision,
|
| 204 |
+
"token": token,
|
| 205 |
+
"cache_dir": cache_dir,
|
| 206 |
+
"local_files_only": local_files_only,
|
| 207 |
+
"trust_remote_code": False,
|
| 208 |
+
}
|
| 209 |
+
)
|
| 210 |
+
backbone = AutoModelForCausalLM.from_pretrained(
|
| 211 |
+
config.base_model_name_or_path,
|
| 212 |
+
**{key: value for key, value in base_kwargs.items() if value is not None},
|
| 213 |
+
)
|
| 214 |
+
model = cls(config, backbone)
|
| 215 |
+
|
| 216 |
+
source = Path(pretrained_model_name_or_path)
|
| 217 |
+
if source.is_dir():
|
| 218 |
+
delta_path = source / config.delta_filename
|
| 219 |
+
else:
|
| 220 |
+
resolved = cached_file(
|
| 221 |
+
str(pretrained_model_name_or_path),
|
| 222 |
+
config.delta_filename,
|
| 223 |
+
revision=revision,
|
| 224 |
+
token=token,
|
| 225 |
+
cache_dir=cache_dir,
|
| 226 |
+
local_files_only=local_files_only,
|
| 227 |
+
)
|
| 228 |
+
if resolved is None:
|
| 229 |
+
raise FileNotFoundError(config.delta_filename)
|
| 230 |
+
delta_path = Path(resolved)
|
| 231 |
+
delta = load_file(str(delta_path), device="cpu")
|
| 232 |
+
current = model.state_dict()
|
| 233 |
+
absent = sorted(set(delta) - set(current))
|
| 234 |
+
mismatched = {
|
| 235 |
+
key: {"delta": tuple(value.shape), "model": tuple(current[key].shape)}
|
| 236 |
+
for key, value in delta.items()
|
| 237 |
+
if key in current and value.shape != current[key].shape
|
| 238 |
+
}
|
| 239 |
+
if absent or mismatched:
|
| 240 |
+
raise RuntimeError(f"Release delta is incompatible: absent={absent}, mismatched={mismatched}")
|
| 241 |
+
model.load_state_dict(delta, strict=False)
|
| 242 |
+
model._release_load_receipt = {
|
| 243 |
+
"delta_file": str(delta_path),
|
| 244 |
+
"tensor_count": len(delta),
|
| 245 |
+
"total_parameters": sum(int(tensor.numel()) for tensor in delta.values()),
|
| 246 |
+
}
|
| 247 |
+
model.eval()
|
| 248 |
+
return model
|
| 249 |
+
|
| 250 |
+
def forward(
|
| 251 |
+
self,
|
| 252 |
+
input_ids: Optional[torch.LongTensor] = None,
|
| 253 |
+
attention_mask: Optional[torch.Tensor] = None,
|
| 254 |
+
position_ids: Optional[torch.LongTensor] = None,
|
| 255 |
+
inputs_embeds: Optional[torch.Tensor] = None,
|
| 256 |
+
labels: Optional[torch.LongTensor] = None,
|
| 257 |
+
max_loops: int = 1,
|
| 258 |
+
loop_selection: int | torch.LongTensor | None = None,
|
| 259 |
+
return_loop_logits: bool = False,
|
| 260 |
+
use_cache: Optional[bool] = None,
|
| 261 |
+
output_attentions: Optional[bool] = None,
|
| 262 |
+
output_hidden_states: Optional[bool] = None,
|
| 263 |
+
return_dict: Optional[bool] = None,
|
| 264 |
+
logits_to_keep: int | torch.Tensor = 0,
|
| 265 |
+
**kwargs: Any,
|
| 266 |
+
) -> RecurrentCausalLMOutput | tuple[Any, ...]:
|
| 267 |
+
if kwargs:
|
| 268 |
+
unknown = ", ".join(sorted(kwargs))
|
| 269 |
+
raise TypeError(f"Unsupported forward keyword(s): {unknown}")
|
| 270 |
+
if int(max_loops) < 1:
|
| 271 |
+
raise ValueError("max_loops must be at least 1")
|
| 272 |
+
if use_cache:
|
| 273 |
+
raise ValueError("KV caching is disabled for the forced-depth recurrent release")
|
| 274 |
+
if output_attentions or output_hidden_states:
|
| 275 |
+
raise ValueError("Attention/hidden-state collection is not exposed by the release loader")
|
| 276 |
+
if input_ids is not None and inputs_embeds is not None:
|
| 277 |
+
raise ValueError("Specify input_ids or inputs_embeds, not both")
|
| 278 |
+
if inputs_embeds is None:
|
| 279 |
+
if input_ids is None:
|
| 280 |
+
raise ValueError("input_ids or inputs_embeds is required")
|
| 281 |
+
inputs_embeds = self.qwen.embed_tokens(input_ids)
|
| 282 |
+
|
| 283 |
+
batch_size, sequence_length = inputs_embeds.shape[:2]
|
| 284 |
+
device = inputs_embeds.device
|
| 285 |
+
if attention_mask is not None:
|
| 286 |
+
attention_mask = attention_mask.to(device=device)
|
| 287 |
+
if position_ids is None:
|
| 288 |
+
position_ids = torch.arange(sequence_length, device=device).unsqueeze(0).expand(batch_size, -1)
|
| 289 |
+
else:
|
| 290 |
+
position_ids = position_ids.to(device=device)
|
| 291 |
+
cache_position = torch.arange(sequence_length, device=device)
|
| 292 |
+
causal_mask = self._causal_mask(attention_mask, inputs_embeds, cache_position)
|
| 293 |
+
position_embeddings = self._rotary_embeddings(inputs_embeds, position_ids)
|
| 294 |
+
|
| 295 |
+
hidden = self._run_layers(
|
| 296 |
+
0,
|
| 297 |
+
self.config.prelude_end,
|
| 298 |
+
inputs_embeds,
|
| 299 |
+
causal_mask,
|
| 300 |
+
position_ids,
|
| 301 |
+
cache_position,
|
| 302 |
+
position_embeddings,
|
| 303 |
+
)
|
| 304 |
+
prelude = hidden
|
| 305 |
+
recurrent_state = hidden
|
| 306 |
+
logits_by_loop: list[torch.Tensor] = []
|
| 307 |
+
for loop_index in range(int(max_loops)):
|
| 308 |
+
loop_input = recurrent_state if loop_index == 0 else self.bridge(recurrent_state, prelude)
|
| 309 |
+
recurrent_state = self._run_layers(
|
| 310 |
+
self.config.prelude_end,
|
| 311 |
+
self.config.recurrent_end,
|
| 312 |
+
loop_input,
|
| 313 |
+
causal_mask,
|
| 314 |
+
position_ids,
|
| 315 |
+
cache_position,
|
| 316 |
+
position_embeddings,
|
| 317 |
+
)
|
| 318 |
+
coda = self._run_layers(
|
| 319 |
+
self.config.recurrent_end,
|
| 320 |
+
len(self.qwen.layers),
|
| 321 |
+
recurrent_state,
|
| 322 |
+
causal_mask,
|
| 323 |
+
position_ids,
|
| 324 |
+
cache_position,
|
| 325 |
+
position_embeddings,
|
| 326 |
+
)
|
| 327 |
+
normed = self.qwen.norm(coda)
|
| 328 |
+
logits_by_loop.append(self.lm_head(self._slice_for_logits(normed, logits_to_keep)))
|
| 329 |
+
|
| 330 |
+
loop_logits = torch.stack(logits_by_loop, dim=1)
|
| 331 |
+
selected_counts = self._resolve_loop_selection(loop_selection, batch_size, int(max_loops), device)
|
| 332 |
+
batch_indices = torch.arange(batch_size, device=device)
|
| 333 |
+
logits = loop_logits[batch_indices, selected_counts - 1]
|
| 334 |
+
loss = None
|
| 335 |
+
if labels is not None:
|
| 336 |
+
if not (isinstance(logits_to_keep, int) and logits_to_keep == 0):
|
| 337 |
+
raise ValueError("labels require logits_to_keep=0")
|
| 338 |
+
shifted_logits = logits[:, :-1, :].contiguous().float()
|
| 339 |
+
shifted_labels = labels[:, 1:].contiguous().to(device=device)
|
| 340 |
+
loss = F.cross_entropy(
|
| 341 |
+
shifted_logits.view(-1, shifted_logits.shape[-1]),
|
| 342 |
+
shifted_labels.view(-1),
|
| 343 |
+
ignore_index=-100,
|
| 344 |
+
)
|
| 345 |
+
|
| 346 |
+
output = RecurrentCausalLMOutput(
|
| 347 |
+
loss=loss,
|
| 348 |
+
logits=logits,
|
| 349 |
+
past_key_values=None,
|
| 350 |
+
hidden_states=None,
|
| 351 |
+
attentions=None,
|
| 352 |
+
loop_logits=loop_logits if return_loop_logits else None,
|
| 353 |
+
selected_loop_counts=selected_counts,
|
| 354 |
+
)
|
| 355 |
+
return output if return_dict is not False else output.to_tuple()
|
| 356 |
+
|
| 357 |
+
def prepare_inputs_for_generation(
|
| 358 |
+
self,
|
| 359 |
+
input_ids: torch.LongTensor,
|
| 360 |
+
attention_mask: Optional[torch.Tensor] = None,
|
| 361 |
+
**kwargs: Any,
|
| 362 |
+
) -> dict[str, Any]:
|
| 363 |
+
return {
|
| 364 |
+
"input_ids": input_ids,
|
| 365 |
+
"attention_mask": attention_mask,
|
| 366 |
+
"max_loops": int(kwargs.get("max_loops", 1)),
|
| 367 |
+
"loop_selection": kwargs.get("loop_selection"),
|
| 368 |
+
"use_cache": False,
|
| 369 |
+
}
|
| 370 |
+
|
| 371 |
+
@staticmethod
|
| 372 |
+
def _resolve_loop_selection(
|
| 373 |
+
selection: int | torch.LongTensor | None,
|
| 374 |
+
batch_size: int,
|
| 375 |
+
max_loops: int,
|
| 376 |
+
device: torch.device,
|
| 377 |
+
) -> torch.LongTensor:
|
| 378 |
+
if selection is None:
|
| 379 |
+
counts = torch.full((batch_size,), max_loops, dtype=torch.long, device=device)
|
| 380 |
+
elif isinstance(selection, int):
|
| 381 |
+
counts = torch.full((batch_size,), int(selection), dtype=torch.long, device=device)
|
| 382 |
+
else:
|
| 383 |
+
counts = selection.to(device=device, dtype=torch.long).reshape(-1)
|
| 384 |
+
if counts.numel() != batch_size:
|
| 385 |
+
raise ValueError("loop_selection tensor must contain one value per batch row")
|
| 386 |
+
if bool(((counts < 1) | (counts > max_loops)).any()):
|
| 387 |
+
raise ValueError("loop_selection values must lie in [1, max_loops]")
|
| 388 |
+
return counts
|
| 389 |
+
|
| 390 |
+
def _causal_mask(
|
| 391 |
+
self,
|
| 392 |
+
attention_mask: Optional[torch.Tensor],
|
| 393 |
+
inputs_embeds: torch.Tensor,
|
| 394 |
+
cache_position: torch.Tensor,
|
| 395 |
+
) -> torch.Tensor | None:
|
| 396 |
+
update = getattr(self.qwen, "_update_causal_mask", None)
|
| 397 |
+
if update is not None:
|
| 398 |
+
return self._call_supported(
|
| 399 |
+
update,
|
| 400 |
+
{
|
| 401 |
+
"attention_mask": attention_mask,
|
| 402 |
+
"input_tensor": inputs_embeds,
|
| 403 |
+
"inputs_embeds": inputs_embeds,
|
| 404 |
+
"cache_position": cache_position,
|
| 405 |
+
"past_key_values": None,
|
| 406 |
+
"output_attentions": False,
|
| 407 |
+
},
|
| 408 |
+
)
|
| 409 |
+
batch_size, sequence_length = inputs_embeds.shape[:2]
|
| 410 |
+
minimum = torch.finfo(inputs_embeds.dtype).min
|
| 411 |
+
causal = torch.full(
|
| 412 |
+
(sequence_length, sequence_length),
|
| 413 |
+
minimum,
|
| 414 |
+
dtype=inputs_embeds.dtype,
|
| 415 |
+
device=inputs_embeds.device,
|
| 416 |
+
).triu(diagonal=1)
|
| 417 |
+
causal = causal.unsqueeze(0).unsqueeze(0).expand(batch_size, 1, -1, -1)
|
| 418 |
+
if attention_mask is not None:
|
| 419 |
+
causal = causal.masked_fill(attention_mask[:, None, None, :].eq(0), minimum)
|
| 420 |
+
return causal
|
| 421 |
+
|
| 422 |
+
def _rotary_embeddings(
|
| 423 |
+
self,
|
| 424 |
+
hidden_states: torch.Tensor,
|
| 425 |
+
position_ids: torch.Tensor,
|
| 426 |
+
) -> Any:
|
| 427 |
+
rotary = getattr(self.qwen, "rotary_emb", None)
|
| 428 |
+
if rotary is None:
|
| 429 |
+
return None
|
| 430 |
+
try:
|
| 431 |
+
return rotary(hidden_states, position_ids)
|
| 432 |
+
except TypeError:
|
| 433 |
+
return None
|
| 434 |
+
|
| 435 |
+
def _run_layers(
|
| 436 |
+
self,
|
| 437 |
+
start: int,
|
| 438 |
+
end: int,
|
| 439 |
+
hidden_states: torch.Tensor,
|
| 440 |
+
causal_mask: Optional[torch.Tensor],
|
| 441 |
+
position_ids: torch.Tensor,
|
| 442 |
+
cache_position: torch.Tensor,
|
| 443 |
+
position_embeddings: Any,
|
| 444 |
+
) -> torch.Tensor:
|
| 445 |
+
for layer in self.qwen.layers[start:end]:
|
| 446 |
+
parameters = inspect.signature(layer.forward).parameters
|
| 447 |
+
cache_key = "past_key_values" if "past_key_values" in parameters else "past_key_value"
|
| 448 |
+
outputs = layer(
|
| 449 |
+
hidden_states,
|
| 450 |
+
**self._filter_supported(
|
| 451 |
+
layer.forward,
|
| 452 |
+
{
|
| 453 |
+
"attention_mask": causal_mask,
|
| 454 |
+
"position_ids": position_ids,
|
| 455 |
+
cache_key: None,
|
| 456 |
+
"output_attentions": False,
|
| 457 |
+
"use_cache": False,
|
| 458 |
+
"cache_position": cache_position,
|
| 459 |
+
"position_embeddings": position_embeddings,
|
| 460 |
+
},
|
| 461 |
+
),
|
| 462 |
+
)
|
| 463 |
+
hidden_states = outputs[0] if isinstance(outputs, tuple) else outputs
|
| 464 |
+
return hidden_states
|
| 465 |
+
|
| 466 |
+
@staticmethod
|
| 467 |
+
def _filter_supported(function: Any, kwargs: dict[str, Any]) -> dict[str, Any]:
|
| 468 |
+
parameters = inspect.signature(function).parameters
|
| 469 |
+
accepts_kwargs = any(value.kind == inspect.Parameter.VAR_KEYWORD for value in parameters.values())
|
| 470 |
+
if accepts_kwargs:
|
| 471 |
+
return {key: value for key, value in kwargs.items() if value is not None}
|
| 472 |
+
return {
|
| 473 |
+
key: value
|
| 474 |
+
for key, value in kwargs.items()
|
| 475 |
+
if key in parameters and (value is not None or key in {"use_cache", "output_attentions"})
|
| 476 |
+
}
|
| 477 |
+
|
| 478 |
+
@classmethod
|
| 479 |
+
def _call_supported(cls, function: Any, kwargs: dict[str, Any]) -> Any:
|
| 480 |
+
return function(**cls._filter_supported(function, kwargs))
|
| 481 |
+
|
| 482 |
+
@staticmethod
|
| 483 |
+
def _slice_for_logits(hidden_states: torch.Tensor, logits_to_keep: int | torch.Tensor) -> torch.Tensor:
|
| 484 |
+
if isinstance(logits_to_keep, int):
|
| 485 |
+
return hidden_states if logits_to_keep == 0 else hidden_states[:, -logits_to_keep:, :]
|
| 486 |
+
return hidden_states[:, logits_to_keep, :]
|
recurrent_delta.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:a15147dc4338c2ba21b3dc4ec440824516344153fe16bd95ac33e3d726c6d1c0
|
| 3 |
+
size 364348444
|
verification_spec.json
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"expected_correct_by_depth": {
|
| 3 |
+
"1": 8
|
| 4 |
+
},
|
| 5 |
+
"source_data": "outputs/stage5/stage5_natural_surface_receipts_20260709_210151/data/robust_relay_fronted_d1_12.jsonl",
|
| 6 |
+
"source_receipt": "outputs/stage5/stage5_natural_surface_followups_2_3_20260710/active/step_2000/step_2000_robust_relay_fronted_d1_12_active_rows_sample.jsonl",
|
| 7 |
+
"identity_check": true,
|
| 8 |
+
"verification_data": "hf_release/verification_assets/recurrent-qwen2.5-0.5b-natural-keeper.jsonl",
|
| 9 |
+
"verification_data_sha256": "81e598aab8cc77cdde5dc5484a66eb2804f3759640ce92244fcf897b648231dd",
|
| 10 |
+
"row_count": 8,
|
| 11 |
+
"rows_by_depth": {
|
| 12 |
+
"1": 8
|
| 13 |
+
}
|
| 14 |
+
}
|
verification_subset.jsonl
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{"answer_text": "Ted", "chain_answer_by_loop": {"1": "Ted"}, "chain_symbol_by_loop": {"1": "Ted"}, "completion": " Ted", "depth": 1, "id": "test_relay_d01_00000", "instance_id": "test_relay_d01_00000", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Ted"], "loop_completions": [" Ted"], "mapping": {"Ada": "Lee", "Amy": "Joe", "Ana": "Ann", "Ann": "Ray", "Ben": "Ada", "Bob": "Bob", "Dan": "Ted", "Jan": "Ada", "Joe": "Max", "Jon": "Bob", "Kim": "Tom", "Lee": "Una", "Max": "Ann", "Ray": "Dan", "Sam": "Ana", "Ted": "Lee", "Tim": "Ben", "Tom": "Jon", "Una": "Ted", "Val": "Kim"}, "n_symbols": 20, "orbit": ["Dan", "Ted"], "prompt_style": "question_only", "question": "Whenever Jon has the key, next it goes to Bob.\nWhenever Dan has the key, next it goes to Ted.\nWhenever Tom has the key, next it goes to Jon.\nWhenever Ann has the key, next it goes to Ray.\nWhenever Ana has the key, next it goes to Ann.\nWhenever Val has the key, next it goes to Kim.\nWhenever Joe has the key, next it goes to Max.\nWhenever Una has the key, next it goes to Ted.\nWhenever Lee has the key, next it goes to Una.\nWhenever Amy has the key, next it goes to Joe.\nWhenever Kim has the key, next it goes to Tom.\nWhenever Tim has the key, next it goes to Ben.\nWhenever Max has the key, next it goes to Ann.\nWhenever Sam has the key, next it goes to Ana.\nWhenever Ben has the key, next it goes to Ada.\nWhenever Ted has the key, next it goes to Lee.\nWhenever Ada has the key, next it goes to Lee.\nWhenever Bob has the key, next it goes to Bob.\nWhenever Jan has the key, next it goes to Ada.\nWhenever Ray has the key, next it goes to Dan.\n\nStarting with Dan, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Dan", "step_sentences": ["After day 1, Ted has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Ted", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 2 |
+
{"answer_text": "Bob", "chain_answer_by_loop": {"1": "Bob"}, "chain_symbol_by_loop": {"1": "Bob"}, "completion": " Bob", "depth": 1, "id": "test_relay_d01_00001", "instance_id": "test_relay_d01_00001", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Bob"], "loop_completions": [" Bob"], "mapping": {"Ada": "Ted", "Amy": "Val", "Ana": "Ann", "Ann": "Sam", "Ben": "Bob", "Bob": "Amy", "Dan": "Sam", "Jan": "Val", "Joe": "Ben", "Jon": "Bob", "Kim": "Joe", "Lee": "Bob", "Max": "Ann", "Ray": "Ted", "Sam": "Val", "Ted": "Ray", "Tim": "Ada", "Tom": "Tim", "Una": "Ana", "Val": "Ana"}, "n_symbols": 20, "orbit": ["Ben", "Bob"], "prompt_style": "question_only", "question": "Whenever Ray has the key, next it goes to Ted.\nWhenever Sam has the key, next it goes to Val.\nWhenever Max has the key, next it goes to Ann.\nWhenever Joe has the key, next it goes to Ben.\nWhenever Amy has the key, next it goes to Val.\nWhenever Ann has the key, next it goes to Sam.\nWhenever Bob has the key, next it goes to Amy.\nWhenever Ada has the key, next it goes to Ted.\nWhenever Tim has the key, next it goes to Ada.\nWhenever Lee has the key, next it goes to Bob.\nWhenever Ben has the key, next it goes to Bob.\nWhenever Dan has the key, next it goes to Sam.\nWhenever Jon has the key, next it goes to Bob.\nWhenever Una has the key, next it goes to Ana.\nWhenever Val has the key, next it goes to Ana.\nWhenever Kim has the key, next it goes to Joe.\nWhenever Ana has the key, next it goes to Ann.\nWhenever Jan has the key, next it goes to Val.\nWhenever Ted has the key, next it goes to Ray.\nWhenever Tom has the key, next it goes to Tim.\n\nStarting with Ben, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Ben", "step_sentences": ["After day 1, Bob has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Bob", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 3 |
+
{"answer_text": "Bob", "chain_answer_by_loop": {"1": "Bob"}, "chain_symbol_by_loop": {"1": "Bob"}, "completion": " Bob", "depth": 1, "id": "test_relay_d01_00002", "instance_id": "test_relay_d01_00002", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Bob"], "loop_completions": [" Bob"], "mapping": {"Ada": "Val", "Amy": "Jon", "Ana": "Ana", "Ann": "Bob", "Ben": "Ben", "Bob": "Tom", "Dan": "Dan", "Jan": "Bob", "Joe": "Ben", "Jon": "Jon", "Kim": "Lee", "Lee": "Amy", "Max": "Joe", "Ray": "Val", "Sam": "Sam", "Ted": "Tim", "Tim": "Joe", "Tom": "Sam", "Una": "Ana", "Val": "Dan"}, "n_symbols": 20, "orbit": ["Jan", "Bob"], "prompt_style": "question_only", "question": "Whenever Joe has the key, next it goes to Ben.\nWhenever Lee has the key, next it goes to Amy.\nWhenever Ben has the key, next it goes to Ben.\nWhenever Jon has the key, next it goes to Jon.\nWhenever Una has the key, next it goes to Ana.\nWhenever Max has the key, next it goes to Joe.\nWhenever Tim has the key, next it goes to Joe.\nWhenever Jan has the key, next it goes to Bob.\nWhenever Kim has the key, next it goes to Lee.\nWhenever Ted has the key, next it goes to Tim.\nWhenever Amy has the key, next it goes to Jon.\nWhenever Ann has the key, next it goes to Bob.\nWhenever Tom has the key, next it goes to Sam.\nWhenever Bob has the key, next it goes to Tom.\nWhenever Sam has the key, next it goes to Sam.\nWhenever Ana has the key, next it goes to Ana.\nWhenever Ray has the key, next it goes to Val.\nWhenever Dan has the key, next it goes to Dan.\nWhenever Ada has the key, next it goes to Val.\nWhenever Val has the key, next it goes to Dan.\n\nStarting with Jan, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Jan", "step_sentences": ["After day 1, Bob has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Bob", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 4 |
+
{"answer_text": "Ted", "chain_answer_by_loop": {"1": "Ted"}, "chain_symbol_by_loop": {"1": "Ted"}, "completion": " Ted", "depth": 1, "id": "test_relay_d01_00003", "instance_id": "test_relay_d01_00003", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Ted"], "loop_completions": [" Ted"], "mapping": {"Ada": "Tim", "Amy": "Ted", "Ana": "Una", "Ann": "Sam", "Ben": "Ted", "Bob": "Ana", "Dan": "Ana", "Jan": "Tom", "Joe": "Tom", "Jon": "Ray", "Kim": "Lee", "Lee": "Una", "Max": "Tom", "Ray": "Dan", "Sam": "Ben", "Ted": "Ann", "Tim": "Bob", "Tom": "Max", "Una": "Jan", "Val": "Lee"}, "n_symbols": 20, "orbit": ["Amy", "Ted"], "prompt_style": "question_only", "question": "Whenever Ann has the key, next it goes to Sam.\nWhenever Dan has the key, next it goes to Ana.\nWhenever Ada has the key, next it goes to Tim.\nWhenever Max has the key, next it goes to Tom.\nWhenever Una has the key, next it goes to Jan.\nWhenever Lee has the key, next it goes to Una.\nWhenever Joe has the key, next it goes to Tom.\nWhenever Jon has the key, next it goes to Ray.\nWhenever Tom has the key, next it goes to Max.\nWhenever Ana has the key, next it goes to Una.\nWhenever Ted has the key, next it goes to Ann.\nWhenever Bob has the key, next it goes to Ana.\nWhenever Val has the key, next it goes to Lee.\nWhenever Tim has the key, next it goes to Bob.\nWhenever Jan has the key, next it goes to Tom.\nWhenever Kim has the key, next it goes to Lee.\nWhenever Ray has the key, next it goes to Dan.\nWhenever Amy has the key, next it goes to Ted.\nWhenever Sam has the key, next it goes to Ben.\nWhenever Ben has the key, next it goes to Ted.\n\nStarting with Amy, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Amy", "step_sentences": ["After day 1, Ted has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Ted", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 5 |
+
{"answer_text": "Ted", "chain_answer_by_loop": {"1": "Ted"}, "chain_symbol_by_loop": {"1": "Ted"}, "completion": " Ted", "depth": 1, "id": "test_relay_d01_00004", "instance_id": "test_relay_d01_00004", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Ted"], "loop_completions": [" Ted"], "mapping": {"Ada": "Ted", "Amy": "Bob", "Ana": "Una", "Ann": "Jon", "Ben": "Dan", "Bob": "Ana", "Dan": "Ray", "Jan": "Ted", "Joe": "Tom", "Jon": "Ana", "Kim": "Max", "Lee": "Sam", "Max": "Sam", "Ray": "Max", "Sam": "Una", "Ted": "Bob", "Tim": "Tom", "Tom": "Jon", "Una": "Joe", "Val": "Lee"}, "n_symbols": 20, "orbit": ["Jan", "Ted"], "prompt_style": "question_only", "question": "Whenever Joe has the key, next it goes to Tom.\nWhenever Jan has the key, next it goes to Ted.\nWhenever Lee has the key, next it goes to Sam.\nWhenever Dan has the key, next it goes to Ray.\nWhenever Ana has the key, next it goes to Una.\nWhenever Una has the key, next it goes to Joe.\nWhenever Ann has the key, next it goes to Jon.\nWhenever Kim has the key, next it goes to Max.\nWhenever Tim has the key, next it goes to Tom.\nWhenever Bob has the key, next it goes to Ana.\nWhenever Tom has the key, next it goes to Jon.\nWhenever Jon has the key, next it goes to Ana.\nWhenever Val has the key, next it goes to Lee.\nWhenever Ben has the key, next it goes to Dan.\nWhenever Ray has the key, next it goes to Max.\nWhenever Amy has the key, next it goes to Bob.\nWhenever Ada has the key, next it goes to Ted.\nWhenever Ted has the key, next it goes to Bob.\nWhenever Max has the key, next it goes to Sam.\nWhenever Sam has the key, next it goes to Una.\n\nStarting with Jan, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Jan", "step_sentences": ["After day 1, Ted has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Ted", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 6 |
+
{"answer_text": "Jan", "chain_answer_by_loop": {"1": "Jan"}, "chain_symbol_by_loop": {"1": "Jan"}, "completion": " Jan", "depth": 1, "id": "test_relay_d01_00005", "instance_id": "test_relay_d01_00005", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Jan"], "loop_completions": [" Jan"], "mapping": {"Ada": "Ana", "Amy": "Lee", "Ana": "Bob", "Ann": "Tom", "Ben": "Dan", "Bob": "Ana", "Dan": "Una", "Jan": "Amy", "Joe": "Kim", "Jon": "Max", "Kim": "Una", "Lee": "Sam", "Max": "Ben", "Ray": "Ted", "Sam": "Joe", "Ted": "Dan", "Tim": "Jan", "Tom": "Tom", "Una": "Dan", "Val": "Joe"}, "n_symbols": 20, "orbit": ["Tim", "Jan"], "prompt_style": "question_only", "question": "Whenever Jon has the key, next it goes to Max.\nWhenever Dan has the key, next it goes to Una.\nWhenever Una has the key, next it goes to Dan.\nWhenever Bob has the key, next it goes to Ana.\nWhenever Ted has the key, next it goes to Dan.\nWhenever Val has the key, next it goes to Joe.\nWhenever Tim has the key, next it goes to Jan.\nWhenever Ray has the key, next it goes to Ted.\nWhenever Max has the key, next it goes to Ben.\nWhenever Ana has the key, next it goes to Bob.\nWhenever Jan has the key, next it goes to Amy.\nWhenever Ada has the key, next it goes to Ana.\nWhenever Sam has the key, next it goes to Joe.\nWhenever Joe has the key, next it goes to Kim.\nWhenever Tom has the key, next it goes to Tom.\nWhenever Amy has the key, next it goes to Lee.\nWhenever Lee has the key, next it goes to Sam.\nWhenever Ben has the key, next it goes to Dan.\nWhenever Kim has the key, next it goes to Una.\nWhenever Ann has the key, next it goes to Tom.\n\nStarting with Tim, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Tim", "step_sentences": ["After day 1, Jan has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Jan", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 7 |
+
{"answer_text": "Kim", "chain_answer_by_loop": {"1": "Kim"}, "chain_symbol_by_loop": {"1": "Kim"}, "completion": " Kim", "depth": 1, "id": "test_relay_d01_00006", "instance_id": "test_relay_d01_00006", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Kim"], "loop_completions": [" Kim"], "mapping": {"Ada": "Ray", "Amy": "Lee", "Ana": "Amy", "Ann": "Dan", "Ben": "Jon", "Bob": "Val", "Dan": "Dan", "Jan": "Sam", "Joe": "Ted", "Jon": "Ben", "Kim": "Max", "Lee": "Dan", "Max": "Tim", "Ray": "Kim", "Sam": "Tom", "Ted": "Ben", "Tim": "Amy", "Tom": "Ben", "Una": "Kim", "Val": "Bob"}, "n_symbols": 20, "orbit": ["Una", "Kim"], "prompt_style": "question_only", "question": "Whenever Ana has the key, next it goes to Amy.\nWhenever Val has the key, next it goes to Bob.\nWhenever Lee has the key, next it goes to Dan.\nWhenever Sam has the key, next it goes to Tom.\nWhenever Kim has the key, next it goes to Max.\nWhenever Max has the key, next it goes to Tim.\nWhenever Jon has the key, next it goes to Ben.\nWhenever Ada has the key, next it goes to Ray.\nWhenever Tom has the key, next it goes to Ben.\nWhenever Jan has the key, next it goes to Sam.\nWhenever Dan has the key, next it goes to Dan.\nWhenever Joe has the key, next it goes to Ted.\nWhenever Bob has the key, next it goes to Val.\nWhenever Amy has the key, next it goes to Lee.\nWhenever Una has the key, next it goes to Kim.\nWhenever Ann has the key, next it goes to Dan.\nWhenever Ted has the key, next it goes to Ben.\nWhenever Ben has the key, next it goes to Jon.\nWhenever Ray has the key, next it goes to Kim.\nWhenever Tim has the key, next it goes to Amy.\n\nStarting with Una, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Una", "step_sentences": ["After day 1, Kim has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Kim", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|
| 8 |
+
{"answer_text": "Dan", "chain_answer_by_loop": {"1": "Dan"}, "chain_symbol_by_loop": {"1": "Dan"}, "completion": " Dan", "depth": 1, "id": "test_relay_d01_00007", "instance_id": "test_relay_d01_00007", "intermediate_chain_supervision": true, "k_star": 1, "latent_targets": ["Dan"], "loop_completions": [" Dan"], "mapping": {"Ada": "Ada", "Amy": "Jan", "Ana": "Ada", "Ann": "Dan", "Ben": "Amy", "Bob": "Kim", "Dan": "Tim", "Jan": "Jon", "Joe": "Kim", "Jon": "Max", "Kim": "Dan", "Lee": "Ben", "Max": "Tom", "Ray": "Amy", "Sam": "Jan", "Ted": "Amy", "Tim": "Lee", "Tom": "Lee", "Una": "Ana", "Val": "Ben"}, "n_symbols": 20, "orbit": ["Ann", "Dan"], "prompt_style": "question_only", "question": "Whenever Val has the key, next it goes to Ben.\nWhenever Ted has the key, next it goes to Amy.\nWhenever Tom has the key, next it goes to Lee.\nWhenever Joe has the key, next it goes to Kim.\nWhenever Kim has the key, next it goes to Dan.\nWhenever Dan has the key, next it goes to Tim.\nWhenever Jon has the key, next it goes to Max.\nWhenever Ada has the key, next it goes to Ada.\nWhenever Bob has the key, next it goes to Kim.\nWhenever Lee has the key, next it goes to Ben.\nWhenever Jan has the key, next it goes to Jon.\nWhenever Una has the key, next it goes to Ana.\nWhenever Amy has the key, next it goes to Jan.\nWhenever Sam has the key, next it goes to Jan.\nWhenever Ann has the key, next it goes to Dan.\nWhenever Ray has the key, next it goes to Amy.\nWhenever Ana has the key, next it goes to Ada.\nWhenever Ben has the key, next it goes to Amy.\nWhenever Max has the key, next it goes to Tom.\nWhenever Tim has the key, next it goes to Lee.\n\nStarting with Ann, pass the key forward once each day. Who has it after exactly 1 days?", "score_target": "full_symbols", "split": "test", "start": "Ann", "step_sentences": ["After day 1, Dan has the key."], "symbol_names": ["Ben", "Sam", "Tom", "Max", "Ada", "Lee", "Ana", "Joe", "Amy", "Dan", "Ray", "Ted", "Una", "Val", "Bob", "Ann", "Tim", "Jan", "Kim", "Jon"], "synthetic_depth": 1, "synthetic_task": "natural_surface_iterated_function", "target": "Dan", "target_loop_count": 1, "template_variant": "fronted", "verbal_surface_family": "relay"}
|