Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

peonist-ai
/
halogen-qwen3.8-flash-next

Text Generation
halogen
halogen-flash
qwen
Mixture of Experts
mixture-of-experts
long-context
quantized
4-bit precision
strix-halo
gfx1151
rocm
amd
radeon
ryzen-ai
local-llm
Model card Files Files and versions
xet
Community
9
New discussion
Resources
  • PR & discussions documentation
  • Code of Conduct
  • Hub documentation

Variable MTP and ngram

1
#9 opened 6 days ago by
davenetdev

Can the model size be further reduced to lower initial loading overhead

4
#8 opened 8 days ago by
yhx03

What an absolute godlike tool. The speed is insane!

❤️ 10
22
#7 opened 11 days ago by
DarkCobalt

Q5 or Q6 quants ?

7
#6 opened 13 days ago by
auf1r2

Donato's video on halogen & qwen3.8 flash

👍❤️ 8
2
#5 opened 15 days ago by
iamtc

Ablated model?

40
#4 opened 17 days ago by
haign

How about a 3-bit quant?

8
#3 opened 18 days ago by
PhilTomson

Engram SSD Offloading

4
#2 opened 21 days ago by
justanonymuser

Very fast.

25
#1 opened 22 days ago by
theTross
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs