Instructions to use taobao-mnn/Qwen3.5-2B-Dflash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use taobao-mnn/Qwen3.5-2B-Dflash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="taobao-mnn/Qwen3.5-2B-Dflash", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("taobao-mnn/Qwen3.5-2B-Dflash", trust_remote_code=True) model = AutoModel.from_pretrained("taobao-mnn/Qwen3.5-2B-Dflash", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use taobao-mnn/Qwen3.5-2B-Dflash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "taobao-mnn/Qwen3.5-2B-Dflash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "taobao-mnn/Qwen3.5-2B-Dflash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/taobao-mnn/Qwen3.5-2B-Dflash
- SGLang
How to use taobao-mnn/Qwen3.5-2B-Dflash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "taobao-mnn/Qwen3.5-2B-Dflash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "taobao-mnn/Qwen3.5-2B-Dflash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "taobao-mnn/Qwen3.5-2B-Dflash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "taobao-mnn/Qwen3.5-2B-Dflash", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use taobao-mnn/Qwen3.5-2B-Dflash with Docker Model Runner:
docker model run hf.co/taobao-mnn/Qwen3.5-2B-Dflash
Qwen3.5-2B DFlash Draft Model
A DFlash speculative-decoding draft model for Qwen3.5-2B .
English
What is DFlash?
DFlash is a lightweight speculative-decoding draft method. Instead of decoding one token at a time, the draft model proposes a whole block of candidate tokens in a single forward pass, conditioned on hidden states captured from several intermediate layers of the target model. The target model then verifies the block in parallel and accepts the longest correct prefix. This lets the system emit multiple tokens per target forward pass while keeping the output identical to greedy target decoding, which speeds up inference when decoding is memory-bandwidth bound.
Model
| Base / target model | Qwen3.5-2B |
| Method | DFlash (speculative decoding) |
| Framework | SpecForge |
| Hardware | 8× AMD MI308X |
| Training | 6 epochs |
| Data | ~800K samples: a mixture of NVIDIA Nemotron Post-Training Dataset V2 (Nathawani et al., 2025) and CodeAlpaca (Chaudhary, 2023), regenerated with the target model |
Files
model.safetensors— draft model weightsconfig.json— draft model config
Benchmark
Eval config
- Backend: sglang
- Decoding: greedy, temperature = 0, max_tokens = 2048
- Metric: accept length (AL) = Σ completion_tokens / Σ verify_ct (tokens emitted per target verify pass)
Per dataset
| Category | Dataset | AL (block=8) | AL (block=16) |
|---|---|---|---|
| MATH | gsm8k | 4.707 | 5.574 |
| math500 | 5.152 | 6.383 | |
| aime25 | 4.956 | 5.932 | |
| CODE | humaneval | 4.609 | 5.415 |
| mbpp | 4.507 | 5.301 | |
| lcb | 4.252 | 4.858 | |
| CHAT | mtbench | 2.861 | 3.185 |
| arenahard | 3.113 | 3.407 | |
| SCIENCE QA | gpqa | 4.034 | 4.519 |
Overall
| block | AL | accept rate |
|---|---|---|
| 8 | 4.163 | 45.21% |
| 16 | 4.781 | 25.23% |
中文
DFlash 是什么?
DFlash 是一种轻量级的投机解码(speculative decoding)草稿方法。草稿模型不再逐 token 解码,而是借助从目标模型多个中间层捕获的隐状态,在一次前向中一次性提出一整块(block)候选 token;随后目标模型并行校验该 block,并接受最长的正确前缀。因此在保持与目标模型贪心解码完全一致输出的前提下,每次目标前向可产出多个 token ,进而在访存瓶颈时提高推理速度。
模型信息
| 基座 / 目标模型 | Qwen3.5-2B |
| 方法 | DFlash(投机解码) |
| 训练框架 | SpecForge |
| 硬件 | 8 卡 AMD MI308X |
| 训练轮数 | 6 个 epoch |
| 训练数据 | 约 80 万条样本:NVIDIA Nemotron Post-Training Dataset V2 (Nathawani et al., 2025) 与 CodeAlpaca (Chaudhary, 2023) 的混合集 ; 使用目标模型重新生成 |
文件说明
model.safetensors— 草稿模型权重config.json— 草稿模型配置
评测结果
评测配置
- 后端: sglang
- 解码: greedy, 温度 = 0, max_tokens = 2048
- 指标: 接受长度 (AL) = Σ completion_tokens / Σ verify_ct(每次目标校验前向产出的 token 数)
分数据集
| 类型 | 数据集 | AL (block=8) | AL (block=16) |
|---|---|---|---|
| 数学 | gsm8k | 4.707 | 5.574 |
| math500 | 5.152 | 6.383 | |
| aime25 | 4.956 | 5.932 | |
| 代码 | humaneval | 4.609 | 5.415 |
| mbpp | 4.507 | 5.301 | |
| lcb | 4.252 | 4.858 | |
| 对话 | mtbench | 2.861 | 3.185 |
| arenahard | 3.113 | 3.407 | |
| 科学问答 | gpqa | 4.034 | 4.519 |
总览
| block | AL | 接受率 |
|---|---|---|
| 8 | 4.163 | 45.21% |
| 16 | 4.781 | 25.23% |
- Downloads last month
- 186