Neuralese COCONUT browser graph
This repository hosts the compact ONNX graph used by the Neuralese realtime
latent instrument. It is derived from the MIT-labelled
ModalityDance/latent-tts-coconut
checkpoint at revision 89501ce8cc7cefd020b6e30c475a6772494de0a4.
The graph implements strict COCONUT recurrence: a decoder final hidden state can be supplied directly as the next input embedding before returning to token mode. It is not an ordinary token model with post-hoc activation sonification.
Artifact
coconut-decoder-fp16-storage.onnx
- 124M-parameter GPT-2 architecture
- 12 layers, 12 heads, hidden size 768
- FP16 weight storage with FP32 inputs, arithmetic, outputs, and KV cache
- SHA-256:
3bc35680260edb1c66276bbd65ae347535ec89ccb80914b469c1d8cf55f400bd - Source graph: 622 MiB FP32
- Compact graph: 311 MiB
Native FP16 arithmetic was tested but rejected because its recurrent trajectory fell below the project's browser fidelity gate. The published graph halves the download while preserving FP32 computation.
Measured browser parity
Against the publisher-faithful PyTorch trajectory for the fixed six-state probe:
- identical generated tokens;
- minimum latent cosine:
0.99999919; - maximum absolute latent error:
1.742e-2; - approximately 29–30 ms per model call after warm-up on the development M4 Pro.
The checkpoint's answer to the probe is incorrect (### 6 for 2 + 2). This is
a checkpoint capability result, not a conversion discrepancy. The artifact is
published for artistic research, browser inference, and reproducibility—not as
a reliable general reasoning model.
Attribution
Model weights originate from ModalityDance/latent-tts-coconut. The recurrent
continuous-thought method is based on COCONUT, Training Large Language Models
to Reason in a Continuous Latent Space (Meta FAIR, 2024).