Instructions to use mavis-ai/Qwen3.6-35B-MoE-Q4T with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Trellis
How to use mavis-ai/Qwen3.6-35B-MoE-Q4T with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
mavis-ai/Qwen3.6-35B-MoE-Q4T
Q4T build of Qwen3.6-35B-A3B for R.E.V.I.S. on Apple silicon, bundled with its speculative-decoding drafters (MTP and DFlash).
mavis-ai claims no ownership of this model. It is a data-free quantized derivative of the official Qwen3.6-35B-A3B checkpoint; all rights in the underlying model remain with its original authors, and it is redistributed under the Apache License 2.0. It is not a fine-tune: no training of any kind was applied. Source revision: 995ad96eacd98c81ed38be0c5b274b04031597b0.
Compatibility. The routed experts are stored as trellis-coded tensors (
.trellis,.suh,.svh) and need a special trellis decoding computation. This model therefore currently does not run onmlx-lm(nor onmlx-vlm,transformersor vLLM). It runs in R.E.V.I.S. v1.3.0 and later, whose built-in inference engine provides the kernels.
R.E.V.I.S.
This model is made for R.E.V.I.S., a 100% local Cognitive OS for Multi-Agentic AI on Mac. This model is supported from R.E.V.I.S. v1.3.0.
- 🌐 Official Website: https://mavis-ai.co.jp/revis/
- 🐦 Follow our updates on X: @mavis_ai_jp
Summary
| Base model | Qwen3.6-35B-A3B (multimodal MoE; 40 layers, hidden size 2048, 256 routed experts with 8 active plus a shared expert, context length 262144) |
| Format | Q4T: K4 routed experts + N6 dense |
| Routed experts | 4-bit trellis-coded (K4) |
| Dense / attention / shared-expert / embedding / LM head | 6-bit affine, group size 64 (N6) |
| Routers and selection layers | BF16 |
| Vision tower | BF16 |
| Calibration data | none (data-free) |
| Weight size | 19.18 GB (decimal), 17.9 GiB |
| Bundled drafters | MTP Draft-Q8 (+ ProposalHead), DFlash-Q8 (1.60 GB together) |
| KV cache (runtime) | 8-bit |
Bundle layout
.
├── README.md, LICENSE, .gitattributes
├── config.json, configuration.json, generation_config.json
├── tokenizer.json, tokenizer_config.json, vocab.json, merges.txt, chat_template.jinja
├── preprocessor_config.json, processor_config.json, video_preprocessor_config.json
├── model-core-00001-of-00001.safetensors 2.91 GB dense, attention, embeddings, routers, vision
├── model-experts-0000{1..8}-of-00008.safetensors 8 x 2.03 GB trellis-coded routed experts
├── model.safetensors.index.json
├── dense_n6_manifest.json, q4t_manifest.json, q4t-split-lineage.json, q4t_sources.json
└── drafts/
├── SHA256SUMS
├── Qwen3.6-35B-Draft-Q8/ MTP drafter, 8-bit affine (g64), 0.91 GB weights
├── Qwen3.6-35B-Draft-Q8-ProposalHead/ proposal head for the MTP drafter, 0.22 GB
└── Qwen3.6-35B-DFlash-Q8/ DFlash block drafter, 8-bit affine (g64), 0.44 GB
The core file and the expert files are split so the runtime can map experts independently; expert tensors are byte-identical to the unsplit source pack (q4t-split-lineage.json). Checksums for the files under drafts/ are in drafts/SHA256SUMS; per-file SHA-256 values for the root files are recorded in q4t_manifest.json.
Quantization (Q4T)
The T family (Q5T, Q4T, Q3T) is a fixed, data-free recipe: one width per module class, chosen by a written rule rather than by looking at evaluation data. The formats differ only in the expert rate K and the dense width N. This one is K4 + N6.
- Routed experts, K4. Each expert matrix is stored as 16x16-tile trellis codes (4 bits per weight, MCG codebook), with a 128-point Hadamard rotation and per-channel sign vectors (
suh,svh). Encoding minimizes rotated-weight MSE; there is no calibration set, Hessian rounding, importance matrix or quantization-aware training. The trellis decoder and its kernels are part of the R.E.V.I.S. inference engine. - Everything else linear, N6. Attention and linear-attention projections, shared expert, token embeddings and LM head are 6-bit affine, group size 64, quantized from the BF16 source with
mx.quantize(round to nearest). - BF16 kept as-is. Router (
mlp.gate),shared_expert_gate, linear-attentionin_proj_a/in_proj_b, norms, and the whole vision tower. Expert selection is a discontinuous top-k decision, so these tensors are never quantized.
| Format | Routed experts | Dense |
|---|---|---|
| Q5T | K5 | N8 (8-bit affine, g64) |
| Q4T | K4 | N6 (6-bit affine, g64) |
| Q3T | K3 | N6 (6-bit affine, g64) |
Benchmarks
The question these benchmarks answer is how well each quantization suppresses degradation relative to the full-precision model (Full, BF16). All arms are scored on the same inputs against the same BF16 reference. The final numbers will come from a single benchmark run and are not available yet; the table below is intentionally empty.
| Metric | Full (BF16) | Q8 | Q6 | Q5 | Q4 | Q5T | Q4T | Q3T |
|---|---|---|---|---|---|---|---|---|
| KLD vs Full (lower is better) | ||||||||
| Top-1 agreement with Full | ||||||||
| Accuracy benchmarks (to be defined) |
Interim figures (to be replaced)
Earlier internal measurement, kept only until the benchmark above is filled in. KL divergence of next-token distributions from the BF16 Qwen3.6-35B-A3B reference, teacher-forced, mean over scored positions (lower is better). Evaluation set: 48 documents, 17,877 scored positions. These are distribution-distance measurements on a small set, not task benchmarks, and they are interim: the benchmark table above supersedes them.
| Arm | KLD vs BF16 | Weight size |
|---|---|---|
| Q8 | 0.0066 | n/a |
| Q5T | 0.0090 | 23.8 GB |
| Q6 | 0.0116 | 29.3 GB |
| Q4T (this repository) | 0.0181 | 19.2 GB |
| Q5 | 0.0248 | 24.9 GB |
| Q3T | 0.0356 | 15.15 GB |
| Q4 | 0.0682 | 20.7 GB |
Runtime
- Runtime: the inference engine built into R.E.V.I.S., a native Swift / C++20 / MLX engine with ahead-of-time compiled Metal kernels for Apple silicon. It has no Python, localhost server or just-in-time Metal in the inference path.
- Speculative decoding: MTP is the default drafter; DFlash is optional. Both are bundled under
drafts/. Each drafter is only valid with this target. TheProposalHeadis pinned to this pack's header digest; on any other target the runtime ignores it and keeps the exact output head. - Losslessness: with greedy decoding, output with a drafter is intended to equal plain decoding of this target, up to the rounding ties that the runtime certifies. Drafters are an execution-speed feature of the runtime and do not change the target's output.
- Chips: the runtime targets M1 to M5; development and verification so far were on M4- and M5-family Macs.
- Sampling defaults in
generation_config.jsonfollow the base model card (temperature 1.0,top_k 20,top_p 0.95). - Speed depends on the runtime and the machine, not on this model file, and is therefore not reported here.
Memory
Weights on disk and in memory: 19.18 GB. The three drafters add 1.60 GB if all are loaded. The KV cache is 8-bit and grows with context; the runtime admits a load against the available memory budget and reports the largest context that fits. Budget for weights plus drafters plus KV and transient buffers when choosing a Mac.
Provenance of the bundled drafters
Everything in this repository ships together, so the origin of every bundled component is listed here. All bundled drafters are derivatives of their upstream releases; mavis-ai claims no ownership of any of them, and each keeps its upstream license (Apache License 2.0).
| Directory | What it is | Original authors and upstream repository | Upstream revision |
|---|---|---|---|
| (root) | the target model, converted to this format | Qwen Team, Qwen/Qwen3.6-35B-A3B |
995ad96eacd98c81ed38be0c5b274b04031597b0 |
drafts/Qwen3.6-35B-Draft-Q8/ |
MTP drafter: the multi-token-prediction module shipped inside Qwen3.6-35B-A3B, extracted as a standalone drafter and quantized to 8-bit affine (group size 64); fc, router, shared-expert gate and norms stay BF16 |
Qwen Team, Qwen/Qwen3.6-35B-A3B |
995ad96eacd98c81ed38be0c5b274b04031597b0 |
drafts/Qwen3.6-35B-Draft-Q8-ProposalHead/ |
proposal head for the MTP drafter, made by mavis-ai (see below) | derived from the BF16 lm_head.weight of Qwen/Qwen3.6-35B-A3B |
995ad96eacd98c81ed38be0c5b274b04031597b0 |
drafts/Qwen3.6-35B-DFlash-Q8/ |
DFlash block-diffusion drafter, 8-bit affine (group size 64); fc and norms stay BF16 |
Z-Lab and Modal (joint retrain), z-lab/Qwen3.6-35B-A3B-DFlash, mirrored as modal-labs/Qwen3.6-35B-A3B-DFlash; DFlash method and code: z-lab/dflash, paper arXiv:2602.06036 |
f181eece646affea2c38b2765f1aaa01a9734ccd |
ProposalHead (derived artifact). proposal_head.safetensors is a 3-bit affine (group size 64) copy of the base model's BF16 output head (lm_head.weight). The runtime uses it only to choose which blocks of 128 token ids the MTP drafter's proposal is searched in; the chosen rows are then rescored through the exact output head of this pack, and the target model verifies every proposal, so it cannot change the output. proposal_head.json pins it to the MTP drafter and to this pack (header digests); on any other target the runtime ignores it and keeps the exact head. No training was involved.
The README.md inside each drafter directory restates its origin and license. Upstream benchmark material is not reproduced in this repository.
License and attribution
mavis-ai claims no ownership of this model. This repository redistributes a quantized derivative of Qwen3.6-35B-A3B, released by the Qwen Team under the Apache License 2.0. A copy of the license is in LICENSE.
Modification notice: the official checkpoint was converted to the Q4T layout described above (K4 trellis-coded routed experts, 6-bit affine dense tensors, BF16 routers and vision tower). No fine-tuning, additional training or architecture change was applied.
The bundled drafters keep their upstream licenses (Apache License 2.0); see Provenance of the bundled drafters.
Citation
Please cite the base model:
@misc{qwen36_35b_a3b,
title = {{Qwen3.6-35B-A3B}: Agentic Coding Power, Now Open to All},
url = {https://qwen.ai/blog?id=qwen3.6-35b-a3b},
author = {{Qwen Team}},
month = {April},
year = {2026}
}
and, if you use the DFlash drafter:
@article{chen2026dflash,
title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
journal = {arXiv preprint arXiv:2602.06036},
year = {2026}
}
- Downloads last month
- 171
Model tree for mavis-ai/Qwen3.6-35B-MoE-Q4T
Base model
Qwen/Qwen3.6-35B-A3B