mavis-ai/Qwen3.6-35B-MoE-Q4T

Q4T build of Qwen3.6-35B-A3B for R.E.V.I.S. on Apple silicon, bundled with its speculative-decoding drafters (MTP and DFlash).

mavis-ai claims no ownership of this model. It is a data-free quantized derivative of the official Qwen3.6-35B-A3B checkpoint; all rights in the underlying model remain with its original authors, and it is redistributed under the Apache License 2.0. It is not a fine-tune: no training of any kind was applied. Source revision: 995ad96eacd98c81ed38be0c5b274b04031597b0.

Compatibility. The routed experts are stored as trellis-coded tensors (.trellis, .suh, .svh) and need a special trellis decoding computation. This model therefore currently does not run on mlx-lm (nor on mlx-vlm, transformers or vLLM). It runs in R.E.V.I.S. v1.3.0 and later, whose built-in inference engine provides the kernels.

R.E.V.I.S.

This model is made for R.E.V.I.S., a 100% local Cognitive OS for Multi-Agentic AI on Mac. This model is supported from R.E.V.I.S. v1.3.0.

Summary

Base model Qwen3.6-35B-A3B (multimodal MoE; 40 layers, hidden size 2048, 256 routed experts with 8 active plus a shared expert, context length 262144)
Format Q4T: K4 routed experts + N6 dense
Routed experts 4-bit trellis-coded (K4)
Dense / attention / shared-expert / embedding / LM head 6-bit affine, group size 64 (N6)
Routers and selection layers BF16
Vision tower BF16
Calibration data none (data-free)
Weight size 19.18 GB (decimal), 17.9 GiB
Bundled drafters MTP Draft-Q8 (+ ProposalHead), DFlash-Q8 (1.60 GB together)
KV cache (runtime) 8-bit

Bundle layout

.
├── README.md, LICENSE, .gitattributes
├── config.json, configuration.json, generation_config.json
├── tokenizer.json, tokenizer_config.json, vocab.json, merges.txt, chat_template.jinja
├── preprocessor_config.json, processor_config.json, video_preprocessor_config.json
├── model-core-00001-of-00001.safetensors        2.91 GB   dense, attention, embeddings, routers, vision
├── model-experts-0000{1..8}-of-00008.safetensors  8 x 2.03 GB   trellis-coded routed experts
├── model.safetensors.index.json
├── dense_n6_manifest.json, q4t_manifest.json, q4t-split-lineage.json, q4t_sources.json
└── drafts/
    ├── SHA256SUMS
    ├── Qwen3.6-35B-Draft-Q8/                MTP drafter, 8-bit affine (g64), 0.91 GB weights
    ├── Qwen3.6-35B-Draft-Q8-ProposalHead/   proposal head for the MTP drafter, 0.22 GB
    └── Qwen3.6-35B-DFlash-Q8/               DFlash block drafter, 8-bit affine (g64), 0.44 GB

The core file and the expert files are split so the runtime can map experts independently; expert tensors are byte-identical to the unsplit source pack (q4t-split-lineage.json). Checksums for the files under drafts/ are in drafts/SHA256SUMS; per-file SHA-256 values for the root files are recorded in q4t_manifest.json.

Quantization (Q4T)

The T family (Q5T, Q4T, Q3T) is a fixed, data-free recipe: one width per module class, chosen by a written rule rather than by looking at evaluation data. The formats differ only in the expert rate K and the dense width N. This one is K4 + N6.

  • Routed experts, K4. Each expert matrix is stored as 16x16-tile trellis codes (4 bits per weight, MCG codebook), with a 128-point Hadamard rotation and per-channel sign vectors (suh, svh). Encoding minimizes rotated-weight MSE; there is no calibration set, Hessian rounding, importance matrix or quantization-aware training. The trellis decoder and its kernels are part of the R.E.V.I.S. inference engine.
  • Everything else linear, N6. Attention and linear-attention projections, shared expert, token embeddings and LM head are 6-bit affine, group size 64, quantized from the BF16 source with mx.quantize (round to nearest).
  • BF16 kept as-is. Router (mlp.gate), shared_expert_gate, linear-attention in_proj_a/in_proj_b, norms, and the whole vision tower. Expert selection is a discontinuous top-k decision, so these tensors are never quantized.
Format Routed experts Dense
Q5T K5 N8 (8-bit affine, g64)
Q4T K4 N6 (6-bit affine, g64)
Q3T K3 N6 (6-bit affine, g64)

Benchmarks

The question these benchmarks answer is how well each quantization suppresses degradation relative to the full-precision model (Full, BF16). All arms are scored on the same inputs against the same BF16 reference. The final numbers will come from a single benchmark run and are not available yet; the table below is intentionally empty.

Metric Full (BF16) Q8 Q6 Q5 Q4 Q5T Q4T Q3T
KLD vs Full (lower is better)
Top-1 agreement with Full
Accuracy benchmarks (to be defined)

Interim figures (to be replaced)

Earlier internal measurement, kept only until the benchmark above is filled in. KL divergence of next-token distributions from the BF16 Qwen3.6-35B-A3B reference, teacher-forced, mean over scored positions (lower is better). Evaluation set: 48 documents, 17,877 scored positions. These are distribution-distance measurements on a small set, not task benchmarks, and they are interim: the benchmark table above supersedes them.

Arm KLD vs BF16 Weight size
Q8 0.0066 n/a
Q5T 0.0090 23.8 GB
Q6 0.0116 29.3 GB
Q4T (this repository) 0.0181 19.2 GB
Q5 0.0248 24.9 GB
Q3T 0.0356 15.15 GB
Q4 0.0682 20.7 GB

Runtime

  • Runtime: the inference engine built into R.E.V.I.S., a native Swift / C++20 / MLX engine with ahead-of-time compiled Metal kernels for Apple silicon. It has no Python, localhost server or just-in-time Metal in the inference path.
  • Speculative decoding: MTP is the default drafter; DFlash is optional. Both are bundled under drafts/. Each drafter is only valid with this target. The ProposalHead is pinned to this pack's header digest; on any other target the runtime ignores it and keeps the exact output head.
  • Losslessness: with greedy decoding, output with a drafter is intended to equal plain decoding of this target, up to the rounding ties that the runtime certifies. Drafters are an execution-speed feature of the runtime and do not change the target's output.
  • Chips: the runtime targets M1 to M5; development and verification so far were on M4- and M5-family Macs.
  • Sampling defaults in generation_config.json follow the base model card (temperature 1.0, top_k 20, top_p 0.95).
  • Speed depends on the runtime and the machine, not on this model file, and is therefore not reported here.

Memory

Weights on disk and in memory: 19.18 GB. The three drafters add 1.60 GB if all are loaded. The KV cache is 8-bit and grows with context; the runtime admits a load against the available memory budget and reports the largest context that fits. Budget for weights plus drafters plus KV and transient buffers when choosing a Mac.

Provenance of the bundled drafters

Everything in this repository ships together, so the origin of every bundled component is listed here. All bundled drafters are derivatives of their upstream releases; mavis-ai claims no ownership of any of them, and each keeps its upstream license (Apache License 2.0).

Directory What it is Original authors and upstream repository Upstream revision
(root) the target model, converted to this format Qwen Team, Qwen/Qwen3.6-35B-A3B 995ad96eacd98c81ed38be0c5b274b04031597b0
drafts/Qwen3.6-35B-Draft-Q8/ MTP drafter: the multi-token-prediction module shipped inside Qwen3.6-35B-A3B, extracted as a standalone drafter and quantized to 8-bit affine (group size 64); fc, router, shared-expert gate and norms stay BF16 Qwen Team, Qwen/Qwen3.6-35B-A3B 995ad96eacd98c81ed38be0c5b274b04031597b0
drafts/Qwen3.6-35B-Draft-Q8-ProposalHead/ proposal head for the MTP drafter, made by mavis-ai (see below) derived from the BF16 lm_head.weight of Qwen/Qwen3.6-35B-A3B 995ad96eacd98c81ed38be0c5b274b04031597b0
drafts/Qwen3.6-35B-DFlash-Q8/ DFlash block-diffusion drafter, 8-bit affine (group size 64); fc and norms stay BF16 Z-Lab and Modal (joint retrain), z-lab/Qwen3.6-35B-A3B-DFlash, mirrored as modal-labs/Qwen3.6-35B-A3B-DFlash; DFlash method and code: z-lab/dflash, paper arXiv:2602.06036 f181eece646affea2c38b2765f1aaa01a9734ccd

ProposalHead (derived artifact). proposal_head.safetensors is a 3-bit affine (group size 64) copy of the base model's BF16 output head (lm_head.weight). The runtime uses it only to choose which blocks of 128 token ids the MTP drafter's proposal is searched in; the chosen rows are then rescored through the exact output head of this pack, and the target model verifies every proposal, so it cannot change the output. proposal_head.json pins it to the MTP drafter and to this pack (header digests); on any other target the runtime ignores it and keeps the exact head. No training was involved.

The README.md inside each drafter directory restates its origin and license. Upstream benchmark material is not reproduced in this repository.

License and attribution

mavis-ai claims no ownership of this model. This repository redistributes a quantized derivative of Qwen3.6-35B-A3B, released by the Qwen Team under the Apache License 2.0. A copy of the license is in LICENSE.

Modification notice: the official checkpoint was converted to the Q4T layout described above (K4 trellis-coded routed experts, 6-bit affine dense tensors, BF16 routers and vision tower). No fine-tuning, additional training or architecture change was applied.

The bundled drafters keep their upstream licenses (Apache License 2.0); see Provenance of the bundled drafters.

Citation

Please cite the base model:

@misc{qwen36_35b_a3b,
    title = {{Qwen3.6-35B-A3B}: Agentic Coding Power, Now Open to All},
    url = {https://qwen.ai/blog?id=qwen3.6-35b-a3b},
    author = {{Qwen Team}},
    month = {April},
    year = {2026}
}

and, if you use the DFlash drafter:

@article{chen2026dflash,
  title   = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author  = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  journal = {arXiv preprint arXiv:2602.06036},
  year    = {2026}
}
Downloads last month
171
Safetensors
Model size
11B params
Tensor type
U32
·
BF16
·
F16
·
U16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mavis-ai/Qwen3.6-35B-MoE-Q4T

Quantized
(881)
this model

Paper for mavis-ai/Qwen3.6-35B-MoE-Q4T