Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

nvidia
/
Qwen3.8-Flash-Next-NVFP4

Image-Text-to-Text
Safetensors
Model Optimizer
qwen4_exp
nvidia
ModelOpt
Qwen3.8
quantized
FP4
fp4
conversational
8-bit precision
modelopt
Model card Files Files and versions
xet
Community
9
New discussion
Resources
  • PR & discussions documentation
  • Code of Conduct
  • Hub documentation

4*5090 run nvidia/Qwen3.8-Flash-Next-NVFP4

#9 opened 17 days ago by
zhnagchenchne

NEXTN draft acceptance collapses to 1.5: mtp.* fp8 experts are built unquantized

1
#8 opened 21 days ago by
abcd245

[FIX] SGLang ple_embedding_dtype fp8

#7 opened 21 days ago by
voves

Upload avatar.png

#6 opened 21 days ago by deleted

Quant requests?

2
#5 opened about 1 month ago by
pathosethoslogos

Doesn't fit Spark

11
#4 opened about 1 month ago by
turowicz

Any person have a typo problem running qwen3.8 flash next on vllm or sglang?

3
#3 opened about 1 month ago by
yimzaza1k

Fails to run on RTX Pro 6000

👍 1
7
#2 opened about 1 month ago by
dicksondickson

Any chance for a GGUF version?

➕ 3
1
#1 opened about 1 month ago by
kexar
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs