Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Duplicated from
RedHatAI/Qwen3.5-397B-A17B-speculator.dflash
inference-optimization
/
Qwen3.5-397B-A17B-speculator.dflash-extra
like
0
Follow
Inference Optimization
107
Safetensors
speculators
speculative-decoding
dflash
custom_code
arxiv:
2602.06036
License:
apache-2.0
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
Qwen3.5-397B-A17B-speculator.dflash-extra
/
mtp_plots
888 kB
Ctrl+K
Ctrl+K
1 contributor
History:
1 commit
ChibuUkachi
add plots
2d663e7
verified
about 1 month ago
speedup_HumanEval_latency.png
Safe
121 kB
xet
add plots
about 1 month ago
speedup_math_reasoning_latency.png
Safe
126 kB
xet
add plots
about 1 month ago
speedup_qa_latency.png
Safe
127 kB
xet
add plots
about 1 month ago
speedup_question_latency.png
Safe
130 kB
xet
add plots
about 1 month ago
speedup_rag_latency.png
Safe
127 kB
xet
add plots
about 1 month ago
speedup_summarization_latency.png
132 kB
xet
add plots
about 1 month ago
speedup_tool_call_latency.png
Safe
124 kB
xet
add plots
about 1 month ago