--- license: mit library_name: onnxruntime tags: - onnx - webgpu - continuous-thought - coconut --- # Neuralese COCONUT browser graph This repository hosts the compact ONNX graph used by the **Neuralese realtime latent instrument**. It is derived from the MIT-labelled [`ModalityDance/latent-tts-coconut`](https://huggingface.co/ModalityDance/latent-tts-coconut) checkpoint at revision `89501ce8cc7cefd020b6e30c475a6772494de0a4`. The graph implements strict COCONUT recurrence: a decoder final hidden state can be supplied directly as the next input embedding before returning to token mode. It is not an ordinary token model with post-hoc activation sonification. ## Artifact `coconut-decoder-fp16-storage.onnx` - 124M-parameter GPT-2 architecture - 12 layers, 12 heads, hidden size 768 - FP16 weight storage with FP32 inputs, arithmetic, outputs, and KV cache - SHA-256: `3bc35680260edb1c66276bbd65ae347535ec89ccb80914b469c1d8cf55f400bd` - Source graph: 622 MiB FP32 - Compact graph: 311 MiB Native FP16 arithmetic was tested but rejected because its recurrent trajectory fell below the project's browser fidelity gate. The published graph halves the download while preserving FP32 computation. ## Measured browser parity Against the publisher-faithful PyTorch trajectory for the fixed six-state probe: - identical generated tokens; - minimum latent cosine: `0.99999919`; - maximum absolute latent error: `1.742e-2`; - approximately 29–30 ms per model call after warm-up on the development M4 Pro. The checkpoint's answer to the probe is incorrect (`### 6` for `2 + 2`). This is a checkpoint capability result, not a conversion discrepancy. The artifact is published for artistic research, browser inference, and reproducibility—not as a reliable general reasoning model. ## Attribution Model weights originate from `ModalityDance/latent-tts-coconut`. The recurrent continuous-thought method is based on COCONUT, *Training Large Language Models to Reason in a Continuous Latent Space* (Meta FAIR, 2024).