馃捇 GitHub  路   馃 Hugging Face  路   馃寪 Website

Nex-N2.5-Pro-DFlash

A DFlash draft model for speculative decoding with Nex-N2.5-Pro, supporting text and multimodal workloads.

This repository contains only the draft model. It is not a standalone language model: it is loaded alongside Nex-N2.5-Pro and proposes a block of tokens per step, which the target model then verifies. The output distribution is that of Nex-N2.5-Pro; the draft model only reduces latency.

Model Details

Target model Nex-N2.5-Pro
Algorithm DFlash
Parameters 1.29B
Precision BF16
Draft block size 16
Modality Text and multimodal (via the target model)

Training

This model was trained with the SpecForge library on Nex-AGI's in-house data, together with a portion of OpenPerfectBlend.

Training uses a 0.1 CE + 0.9 TV loss, temperature 0.7, a DFlash block size of 16, and 512 anchors per sample.

Deployment

Serve Nex-N2.5-Pro with the draft model attached using sglang. Speculative decoding with this draft model requires our customized sglang fork, which is preinstalled in the prebuilt Docker image nexagi/sglang:v0.5.18-nex-patch:

python -m sglang.launch_server \
  --model-path <NEX_N2_5_PRO_MODEL_PATH> \
  --chat-template <NEX_N2_5_PRO_CHAT_TEMPLATE_PATH> \
  --speculative-draft-model-path <NEX_N2_5_PRO_DFLASH_MODEL_PATH> \
  --speculative-algorithm DFLASH \
  --speculative-dflash-block-size 16 \
  --tp-size 8 \
  --context-length 262144 \
  --max-total-tokens 262144 \
  --mem-fraction-static 0.80 \
  --chunked-prefill-size 8192 \
  --host 0.0.0.0 \
  --port 30000

For the target model's own deployment options, sampling parameters, and thinking modes, see the Nex-N2.5-Pro model card.

Downloads last month
44
Safetensors
Model size
1B params
Tensor type
BF16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for nex-agi/Nex-N2.5-Pro-DFlash

Finetuned
(1)
this model