Qwen3.5-2B DFlash Draft Model

A DFlash speculative-decoding draft model for Qwen3.5-2B .

English

What is DFlash?

DFlash is a lightweight speculative-decoding draft method. Instead of decoding one token at a time, the draft model proposes a whole block of candidate tokens in a single forward pass, conditioned on hidden states captured from several intermediate layers of the target model. The target model then verifies the block in parallel and accepts the longest correct prefix. This lets the system emit multiple tokens per target forward pass while keeping the output identical to greedy target decoding, which speeds up inference when decoding is memory-bandwidth bound.

Model

Base / target model Qwen3.5-2B
Method DFlash (speculative decoding)
Framework SpecForge
Hardware 8× AMD MI308X
Training 6 epochs
Data ~800K samples: a mixture of NVIDIA Nemotron Post-Training Dataset V2 (Nathawani et al., 2025) and CodeAlpaca (Chaudhary, 2023), regenerated with the target model

Files

  • model.safetensors — draft model weights
  • config.json — draft model config

Benchmark

Eval config

  • Backend: sglang
  • Decoding: greedy, temperature = 0, max_tokens = 2048
  • Metric: accept length (AL) = Σ completion_tokens / Σ verify_ct (tokens emitted per target verify pass)

Per dataset

Category Dataset AL (block=8) AL (block=16)
MATH gsm8k 4.707 5.574
math500 5.152 6.383
aime25 4.956 5.932
CODE humaneval 4.609 5.415
mbpp 4.507 5.301
lcb 4.252 4.858
CHAT mtbench 2.861 3.185
arenahard 3.113 3.407
SCIENCE QA gpqa 4.034 4.519

Overall

block AL accept rate
8 4.163 45.21%
16 4.781 25.23%
中文

DFlash 是什么?

DFlash 是一种轻量级的投机解码(speculative decoding)草稿方法。草稿模型不再逐 token 解码,而是借助从目标模型多个中间层捕获的隐状态,在一次前向中一次性提出一整块(block)候选 token;随后目标模型并行校验该 block,并接受最长的正确前缀。因此在保持与目标模型贪心解码完全一致输出的前提下,每次目标前向可产出多个 token ,进而在访存瓶颈时提高推理速度。

模型信息

基座 / 目标模型 Qwen3.5-2B
方法 DFlash(投机解码)
训练框架 SpecForge
硬件 8 卡 AMD MI308X
训练轮数 6 个 epoch
训练数据 约 80 万条样本:NVIDIA Nemotron Post-Training Dataset V2 (Nathawani et al., 2025) 与 CodeAlpaca (Chaudhary, 2023) 的混合集 ; 使用目标模型重新生成

文件说明

  • model.safetensors — 草稿模型权重
  • config.json — 草稿模型配置

评测结果

评测配置

  • 后端: sglang
  • 解码: greedy, 温度 = 0, max_tokens = 2048
  • 指标: 接受长度 (AL) = Σ completion_tokens / Σ verify_ct(每次目标校验前向产出的 token 数)

分数据集

类型 数据集 AL (block=8) AL (block=16)
数学 gsm8k 4.707 5.574
math500 5.152 6.383
aime25 4.956 5.932
代码 humaneval 4.609 5.415
mbpp 4.507 5.301
lcb 4.252 4.858
对话 mtbench 2.861 3.185
arenahard 3.113 3.407
科学问答 gpqa 4.034 4.519

总览

block AL 接受率
8 4.163 45.21%
16 4.781 25.23%
Downloads last month
186
Safetensors
Model size
0.3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including taobao-mnn/Qwen3.5-2B-Dflash