Neuralese COCONUT browser graph

This repository hosts the compact ONNX graph used by the Neuralese realtime latent instrument. It is derived from the MIT-labelled ModalityDance/latent-tts-coconut checkpoint at revision 89501ce8cc7cefd020b6e30c475a6772494de0a4.

The graph implements strict COCONUT recurrence: a decoder final hidden state can be supplied directly as the next input embedding before returning to token mode. It is not an ordinary token model with post-hoc activation sonification.

Artifact

coconut-decoder-fp16-storage.onnx

  • 124M-parameter GPT-2 architecture
  • 12 layers, 12 heads, hidden size 768
  • FP16 weight storage with FP32 inputs, arithmetic, outputs, and KV cache
  • SHA-256: 3bc35680260edb1c66276bbd65ae347535ec89ccb80914b469c1d8cf55f400bd
  • Source graph: 622 MiB FP32
  • Compact graph: 311 MiB

Native FP16 arithmetic was tested but rejected because its recurrent trajectory fell below the project's browser fidelity gate. The published graph halves the download while preserving FP32 computation.

Measured browser parity

Against the publisher-faithful PyTorch trajectory for the fixed six-state probe:

  • identical generated tokens;
  • minimum latent cosine: 0.99999919;
  • maximum absolute latent error: 1.742e-2;
  • approximately 29–30 ms per model call after warm-up on the development M4 Pro.

The checkpoint's answer to the probe is incorrect (### 6 for 2 + 2). This is a checkpoint capability result, not a conversion discrepancy. The artifact is published for artistic research, browser inference, and reproducibility—not as a reliable general reasoning model.

Attribution

Model weights originate from ModalityDance/latent-tts-coconut. The recurrent continuous-thought method is based on COCONUT, Training Large Language Models to Reason in a Continuous Latent Space (Meta FAIR, 2024).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support