Exactly what I was hoping for, Thx
much appreciated
Amazing, Thank you, this is our record time for comments on an active repo
Amazing, Thank you, this is our record time for comments on an active repo
I was going thru David's model card hoping to find exactly this and there you go, uploaded one minute ago! what are the odds
Hey
We are in the process of publishing a NYaRN version of it, with a built-in YaRN-extended 1 million token context window
Check back in ~15 minutes
Would you care to follow us on platforms for early access to our models?
https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-MXFP4-1M
Here, it behaves as a 1M context model out of the box, with Anvil or similar Turboquant supporting runtimes, the kv cache (~1-2% loos max) ~4 bit will be ~10-20 gb
https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-1M
This is the Nvidia optimised version
Very keen to test that out, its bit late where I am so will get the downloads started first and test them out tmr. Thanks for the good work.
oh and any chance adding mtp version for them as well?
Absolutely, thank you.
Oh yes
Thats one of our main priorities
apart from which, we are also adding mlx ones
oh and any chance adding mtp version for them as well?
will do
Hi, in the model card you still suggest to use Anvil to run this model, but no GGUF is been published in the Files and Versions, so how do you consider to use Anvil?
Thanks
You may refer to our ultraoptimised gguf version, not this model, this model you may wish to run with other inference engines
Here, https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised similar quants will perform similarly, and we are sorry for the inconvenience, this readme was made agentically
I would change the description in the Model Card then... Also, I would remove the GGUF tag from this repo. Regards
Yes thank you again