Instructions to use Bochkov/modular-reasoning-0p5b-g6p5-demo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Bochkov/modular-reasoning-0p5b-g6p5-demo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Bochkov/modular-reasoning-0p5b-g6p5-demo", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Bochkov/modular-reasoning-0p5b-g6p5-demo", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Bochkov/modular-reasoning-0p5b-g6p5-demo with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Bochkov/modular-reasoning-0p5b-g6p5-demo" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Bochkov/modular-reasoning-0p5b-g6p5-demo", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Bochkov/modular-reasoning-0p5b-g6p5-demo
- SGLang
How to use Bochkov/modular-reasoning-0p5b-g6p5-demo with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Bochkov/modular-reasoning-0p5b-g6p5-demo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Bochkov/modular-reasoning-0p5b-g6p5-demo", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Bochkov/modular-reasoning-0p5b-g6p5-demo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Bochkov/modular-reasoning-0p5b-g6p5-demo", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Bochkov/modular-reasoning-0p5b-g6p5-demo with Docker Model Runner:
docker model run hf.co/Bochkov/modular-reasoning-0p5b-g6p5-demo
Modular Residual Reasoning Core with G6.5 Attachment Demo
A research checkpoint of an attention-free reasoning core that integrates typed external elements through commutative residual assembly.
External element types developed in the associated experiments include:
- frozen reconstructive document memory;
- deterministic bounded virtual-machine execution;
- positional tokenizer-aware result buffers.
Checkpoint
| Field | Value |
|---|---|
| Optimizer step | 141092 |
| Processed targets | 4047453941 |
| Stored tensor values | 505,353,758 |
| Approximate parameter values | 504,567,326 |
| Fixed-buffer values | 786,432 |
| Weight SHA-256 | 3c64ccd295b109098923e62edc12f4f30776deaf51b69cad9a6d3543c4863fc8 |
Important distinction
When loaded as a standard causal LM, this repository runs the reasoning core without external memory or VM elements.
The G6.5 attachment result requires the typed external-element protocol
demonstrated in attachment_demo/.
Therefore standard HellaSwag/MMLU/WikiText evaluation and the attachment demo measure different operating modes.
Loading the core
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
path = "Bochkov/modular-reasoning-0p5b-g6p5-demo"
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForCausalLM.from_pretrained(
path,
trust_remote_code=True,
dtype=torch.bfloat16,
).to("cuda").eval()
inputs = tokenizer(
"The capital of France is",
return_tensors="pt",
add_special_tokens=False,
).to("cuda")
with torch.inference_mode():
output = model(
input_ids=inputs["input_ids"],
num_iterations=5,
return_dict=True,
)
print(output.logits.shape)
External consistency assembly
For proposal $z_e$ attached to reasoning node $i$:
Element order is discarded by segment summation.
A deterministic consistency step moves the reasoning field toward the weighted proposal mean before a bounded trainable correction.
Zero-shot G1.7 β G6.5 attachment result
A fixed reasoning checkpoint was evaluated on a provenance-controlled delta suite whose required elements were absent from G1.7 and present in G6.5.
Best reported checkpoint:
| Store | Coverage | Loss | Token accuracy | Gain | Greedy exact |
|---|---|---|---|---|---|
| G1.7 | 0.0000 | 6.5375 | 0.1198 | 0.0000 | 0.0000 |
| G6.5 | 1.0000 | 2.4611 | 0.6381 | 4.0764 nat | 0.3125 |
The reasoning weights were unchanged during attachment.
Fault controls detached 10 required elements. Reattachment restored the baseline exactly under the deterministic evaluation protocol:
reattach_loss_difference = 0
Scope
This is not open-domain QA. The suite contains 417
provenance-controlled tasks referencing 105 newly attached elements.
The experiment measures compatibility, coverage growth and fault
locality.
Demo artifacts
attachment_demo/
βββ qa/
βββ cache/
βββ growth_results.json
βββ run_attachment_demo.py
Run:
python3.11 attachment_demo/run_attachment_demo.py --model . --device cuda:0 --iterations 5
Limitations
- The core was trained on a specialized mixture of language, memory-QA, and VM tasks.
- Standard benchmark scores do not include automatic semantic memory retrieval.
- The demonstrated memory is reconstructive rather than lossless.
- The G6.5 result uses provenance-controlled retrieval tasks.
- No claim is made that stored memory values equal dense trainable parameters.
- No KV cache is implemented.
- Use batch size
1for generic harness evaluation.
Audited standard language-model evaluation β core only
| Benchmark metric | Audited result |
|---|---|
| HellaSwag acc | 26.26 Β± 0.44 |
| HellaSwag acc_norm | 25.31 Β± 0.43 |
| ARC-Easy acc | 28.24 Β± 0.92 |
| ARC-Easy acc_norm | 30.13 Β± 0.94 |
| ARC-Challenge acc | 19.45 Β± 1.16 |
| ARC-Challenge acc_norm | 23.12 Β± 1.23 |
| PIQA acc | 54.03 Β± 1.16 |
| PIQA acc_norm | 51.74 Β± 1.17 |
| WinoGrande acc | 50.20 Β± 1.41 |
| OpenBookQA acc | 12.60 Β± 1.49 |
| OpenBookQA acc_norm | 23.80 Β± 1.91 |
| CommonsenseQA acc | 19.57 Β± 1.14 |
| MMLU 0-shot | 22.95 Β± 0.35 |
| MMLU 5-shot | 22.95 Β± 0.35 |
| LAMBADA accuracy | 0.00 Β± 0.00 |
| LAMBADA perplexity | 1501146.58 Β± 121571.30 |
| WikiText word perplexity | 4006.38 |
| WikiText byte perplexity | 4.72 |
| WikiText bits/byte | 2.24 |
No external memory or VM is enabled in this table. The low standalone LAMBADA/WikiText performance is material to interpreting this artifact; it is not an attachment result.
Bounded attachment demonstration β not a standard LM benchmark
The names G1p7 and G6p5 identify the source configurations used to select evidence. They do not state that billions of memory values are loaded by this downloadable demonstration.
Evaluated: 417 examples, 2396 target tokens. Cache ID: da0ab720f71cc2eb103983b2ae3cc439b354876b85c566406192a6bc5bb796fc.
| Source configuration | Coverage | Target-token NLL | Token accuracy |
|---|---|---|---|
G1p7 |
0.0000 | 6.5716 | 0.1210 |
G6p5 |
1.0000 | 2.8073 | 0.5872 |
Observed NLL difference: 3.7643 nats/token. Because coverage changes from one condition to the other, this difference is not an isolated estimate of memory-scale benefit at fixed answerability.
The physical size of the delivered attachment artifact was not supplied to this update script; no physical-capacity claim is made here.
This is a selected-evidence demonstration. Report the number of selected records, the file size, and the full snapshot footprint separately after auditing them. The source snapshot is not bundled merely because its name appears here.
The QA tasks have controlled provenance. Some candidate-generation experiments used source-evidence token spans and answer offsets provided by the dataset builder; they must not be interpreted as retrieval from an arbitrary user question without those fields.
π§βπ¬ Citation & Concept
If you use this model or the underlying concepts in your research, please cite our work:
@misc{bochkov2026parametermonolithreconstructivememories,
title={Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models},
author={A. Bochkov},
year={2026},
eprint={2610.04012},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.04012},
}
- Downloads last month
- 694