Qwen3.5 ANE Core ML

Run Qwen3.5 text models on Apple Silicon with Core ML. These packages use 6-bit compressed linear weights and include a four-token prompt-prefill graph alongside the one-token decoder.

The repository contains nine model/context packages: 0.8B at 2,048 tokens; 2B at 2,048 and 16,384 tokens; and 4B and 9B at 2,048, 8,192, and 16,384 tokens. This is an unofficial conversion and is not affiliated with or endorsed by the Qwen team. It is based on the original Qwen3.5-0.8B, Qwen3.5-2B, Qwen3.5-4B, and Qwen3.5-9B models.

Highlights

  • Designed for Apple Neural Engine execution. Core ML packages split the model into chunks and keep the shared token embedding table on the host.
  • Compact weights. Linear layers use 6-bit k-means palettization with a group size of 32. Embeddings use row-wise symmetric int8 quantization.
  • Prompt prefill option. Each package contains a four-token graph for processing prompt tokens in four-token groups. Any remaining prompt tokens and all generated tokens use the one-token decoder.
  • Measured token agreement. In short teacher-forced checks, the 2B package matched the original Qwen3.5-2B BF16 model on 60 of 64 tokens (93.75%); the 9B package matched Qwen3.5-9B on 58 of 64 (90.63%).

These agreement scores compare next-token choices on fixed sequences. They are conversion checks, not measures of general task accuracy.

Available artifacts

Model Context Multifunction artifact Core ML size Shared embedding
0.8B 2,048 qwen3.5-0.8b_ctx2048_chunks6_multifunction_lut6_kvslice_gqagrouped 0.54 GiB 243 MiB
2B 2,048 qwen3.5-2b_ctx2048_chunks6_multifunction_lut6_kvslice_gqagrouped 1.33 GiB 485 MiB
2B 16,384 qwen3.5-2b_ctx16384_chunks6_multifunction_lut6_kvslice_gqagrouped 1.35 GiB 485 MiB
4B 2,048 qwen3.5-4b_ctx2048_chunks8_multifunction_lut6_kvslice_gqagrouped 2.98 GiB 607 MiB
4B 8,192 qwen3.5-4b_ctx8192_chunks8_multifunction_lut6_kvslice_gqagrouped 2.99 GiB 607 MiB
4B 16,384 qwen3.5-4b_ctx16384_chunks8_multifunction_lut6_kvslice_gqagrouped 3.00 GiB 607 MiB
9B 2,048 qwen3.5-9b_ctx2048_chunks8_multifunction_lut6_kvslice_gqagrouped 5.59 GiB 970 MiB
9B 8,192 qwen3.5-9b_ctx8192_chunks8_multifunction_lut6_kvslice_gqagrouped 5.60 GiB 970 MiB
9B 16,384 qwen3.5-9b_ctx16384_chunks8_multifunction_lut6_kvslice_gqagrouped 5.61 GiB 970 MiB

Download the selected artifact and the embedding table for its model size. The embedding table is shared across that model's context variants. This example downloads the 2B package with a 2,048-token context and the Python driver used below:

hf download takanori-ishikawa/Qwen3.5-ANE-CoreML \
  --include 'qwen3.5-2b_embed_tokens.int8.bin' \
  --include 'qwen3.5-2b_ctx2048_chunks6_multifunction_lut6_kvslice_gqagrouped/**' \
  --include 'examples/**' \
  --local-dir artifacts

Each package directory contains metadata.json and the .mlpackage chunks for that model and context. The metadata lists the chunks in execution order. On the Hub, each model/context package is a top-level directory, while its shared embedding table is stored at the repository root. This layout matches the paths recorded in metadata.json and lets a runtime select a package by pointing to its directory.

Four-token prefill

Every published artifact contains the standard one-token decoder and a four-token prefill graph. A compatible runtime can process each complete group of four input prompt tokens with the prefill graph. Any remaining prompt tokens and all generated tokens use the one-token decoder. All nine packages listed above include both graphs.

Some older packages predate conversion-signature metadata and mark their compatibility records as legacy_unverified. The 4B/8K, 4B/16K, 9B/8K, and 9B/16K packages include conversion-signature metadata. For 0.8B/2K, five runs confirmed that the prefill path produced the same token IDs as the one-token decoder.

Fidelity

The checks below compare each Core ML-generated token with the top-1 token from the original Hugging Face BF16 model when teacher-forced over the same sequence.

Model Sequence Top-1 agreement Result
2B one-token decode, LUT6 64 tokens 60 / 64 (93.75%) pass
9B one-token decode, LUT6 64 tokens 58 / 64 (90.63%) pass

The 9B graph includes the numerically stable DeltaNet softplus conversion. On a separate fixed 138-token regression sequence, that change improved top-1 agreement from 111/138 (80.43%) to 127/138 (92.03%).

Quantization and Core ML execution can still change generated sequences, especially where the original model has near-tied logits. Validate the model on your own prompts and target OS/device.

Eight of the nine packages passed an end-to-end Swift/Core ML smoke test: the 0.8B/2K package, 2B/2K package, and all 4B and 9B packages. The test loaded each package and generated Qwen ANE ready. The 2B/16K package was not included. This checks integration, not task quality. Longer context changes the static state capacity, but no separate fidelity comparison is available for the longer-context packages.

Runtime status

examples/generate.py is a minimal Python reference driver built with coremltools and NumPy. It supports greedy decoding and configurable temperature, top-k, top-p, and repetition-penalty sampling, with a reproducible random seed. The example shows how to load packages, manage model state, look up host-side embeddings, and run the chunks.

Source code for a Swift runtime with chat templates, sampling, streaming, reset, multi-turn state reuse, and optional four-token prompt prefill is available in the local-llm repository.

Python example

Requirements: an Apple Silicon Mac, Python 3.12, and coremltools 9.

pip install "coremltools>=9,<10" numpy transformers

python artifacts/examples/generate.py \
  --artifact-dir artifacts/qwen3.5-2b_ctx2048_chunks6_multifunction_lut6_kvslice_gqagrouped \
  --prompt "What is the capital of Japan?" \
  --max-new-tokens 64 \
  --temperature 0.7 \
  --top-p 0.8 \
  --top-k 20 \
  --repetition-penalty 1.05 \
  --seed 42

The example loads its tokenizer and chat template from the original Qwen repository. By default, --tokenizer is Qwen/Qwen3.5-2B; set this option to the matching model repository when using another model size. The default --temperature 0 preserves greedy decoding; pass a positive temperature to enable sampling.

Package design

Each artifact contains one metadata.json and a sequence of .mlpackage directories. Run the chunks in metadata order for every token:

  1. The host restores one token embedding from the shared int8 table.
  2. Chunk 0 receives the fp16 embedding and current_pos.
  3. Intermediate chunks exchange fp32 hidden states.
  4. Every chunk updates its Core ML K/V and recurrent state.
  5. The final chunk returns fp32 logits over the full vocabulary.

metadata.json records graph revision, tensor names, chunk order, state layout, embedding checksum, quantization format, and conversion signature. The 9B final vocabulary projection is split into four outputs to keep the last Core ML package below the package-size limit.

Key format details:

  • linear weights: LUT6, k-means, per grouped channel, group size 32
  • embeddings: row-wise symmetric int8 values plus one fp16 scale per row
  • attention cache: split K/V Core ML state with grouped GQA updates
  • DeltaNet: recurrent and convolution state stored in Core ML state
  • position input: graph-generated RoPE and causal mask from current_pos

Observed performance

Measured on a MacBook Pro (Mac14,7, Apple M2, 16 GB) running macOS 15.7.9. Each variant used the same fixed 128-token prompt ID array, greedy decoding, seed 1, two warmup generations, and five measured generations of 64 tokens each. The multifunction artifact's four-token graph handled prompt prefill. Prefill and decode rates are medians across the five runs. Cold load used an empty compiled-model cache; cached load reused that cache in a new process. Each load duration is one observation. Results vary with hardware, OS, cache, context length, and system load.

Model Context Cold load Cached load Prefill median Decode median
0.8B 2,048 49.96 s 1.85 s 127.6 tok/s 50.1 tok/s
2B 2,048 76.77 s 1.96 s 75.4 tok/s 26.0 tok/s
4B 2,048 159.33 s 2.28 s 39.4 tok/s 11.5 tok/s
4B 8,192 163.97 s 2.38 s 19.9 tok/s 10.1 tok/s
4B 16,384 161.41 s 2.38 s 14.3 tok/s 6.5 tok/s
9B 2,048 355.09 s 3.21 s 27.9 tok/s 8.2 tok/s
9B 8,192 343.69 s 5.45 s 21.5 tok/s 7.4 tok/s
9B 16,384 358.66 s 5.57 s 11.7 tok/s 5.1 tok/s

Prompt ID SHA-256: e263916a82c9374465881fbc50252f47a72b46f40ccfe27438d7bd1244e96e43.

Limitations

  • Apple Silicon and Apple Neural Engine execution are required. CPU-only execution is not supported by these packages.
  • Context capacity is static. Choose the 2,048-, 8,192-, or 16,384-token artifact before loading; a loaded package cannot be resized at runtime.
  • Only 4B and 9B are published with 8,192- and 16,384-token contexts. The 0.8B release is limited to 2,048 tokens; 2B is available at 2,048 and 16,384.
  • The Python driver is a minimal text-generation example, not a complete chat application. Swift runtime source is maintained in the linked repository.
  • Top-1 agreement is not a claim of benchmark accuracy or exact output identity.

Source and license

Conversion code and technical notes are maintained in ishikawa/local-llm.

Apache 2.0, the same as the original Qwen3.5 models. The license text is included in LICENSE; the original models are by the Qwen team at Alibaba Cloud.

Downloads last month
1,435
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for takanori-ishikawa/Qwen3.5-ANE-CoreML

Quantized
(316)
this model