Instructions to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: llama cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: llama cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Use Docker
docker model run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- LM Studio
- Jan
- vLLM
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- Ollama
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Ollama:
ollama run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- Unsloth Desktop
- Pi
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Docker Model Runner:
docker model run hf.co/jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
- Lemonade
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Run and chat with the model
lemonade run user.Qwen3.8-Flash-CIRU-STRIX-IU4-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8 Flash CIRU Strix IU4 โ v5.0.1 vision update
Compact weights and a native CIRU runtime for AMD Strix Halo. The original
v4.4.1 GGUF remains in this repository at its existing filename. v5.0 is
an additional, versioned GGUF. Both use the same multi-token prediction
(MTP) draft model and the required SSD-paged embedding files under
ple/.
This package is for Flash. Orca uses a separate model package and its own
matching draft; do not interchange their model or draft files.
This is a custom-runtime model with text and image input. Stock llama.cpp and hosted inference do not support its compact expert format or required SSD-paged embedding table. Build the included matching CIRU runtime and keep all three
ple/files on fast local storage.
Vision in v5.0.1
The v5.0.1 runtime enables image input by default with the matching
projector and keeps MTP depth 3. It restores the QSA attention
position-gap correction needed when image tokens are followed by
speculative decoding. Model, draft and projector weights are unchanged.
Install the new source package and rebuild the runtime; replacing only
the launcher does not apply this fix. Existing v5 users can apply the
small source update without
rebuilding GPU kernels. Set ENABLE_VISION=0 for a
text-only launch. See vision setup and validation.
On Radeon 8060S, Flash IU4 passed six functional checks with its default MTP depth 3: two separate image descriptions, image follow-up, two images in one request, a thinking-mode image request, and a text regression. These are synthetic-image smoke tests, not a broad vision benchmark. Validation record. The speed and memory tables below remain the v5.0 text-only measurements. Loading the vision projector adds encoder memory; those tables do not measure the new image-enabled default.
What changed in v5.0
The routed expert weights use a 4.5175-bits-per-weight compact indexed affine representation. Its GGUF is 72,112,333,312 bytes (67.16 GiB), 6.785 GiB smaller than the retained v4.4.1 GGUF. The runtime uses native packed IU4 for short expert operations and tiled BF16 decoding for matrix prefill. Projection, recurrent-state, SSD paging and long-context attention paths were also improved. IU4 here means four-bit integer dot-product arithmetic on the AMD GPU; it does not mean every model weight uses four-bit storage. These are complete-model gains: the large prefill improvements are not established as IU4-specific. The limits of the same-value IU8 comparison are explained in BENCHMARKS.md.
How the model works
Qwen3.8 Flash is a mixture-of-experts model: each token activates a selected subset of expert networks. This release adapts its storage and inference kernels for 128-GiB AMD Strix Halo systems. It is intended for local chat, coding and tool-enabled applications. It is a quantized adaptation of the base model with a matching native runtime.
The v5.0 expert representation stores four-bit weight codes plus compact indices into scale/offset tables and per-row multipliers. Including those extra values, the expert bank averages 4.5175 bits per weight. The rest of the model keeps mixed higher-precision formats; the entire model is not uniformly four-bit.
For short expert operations, the runtime uses packed four-bit integer dot products, activation scaling and sums, and the matching affine corrections. For long prompt batches, it reconstructs only the temporary BF16 tiles needed by the GPU and reuses efficient batched expert matrix kernels. The resident expert weights remain compact. These two paths let the runtime use appropriate arithmetic for different workloads.
A separate large n-gram embedding table is paged from fast local NVMe
through the three required ple/ files. It is distinct from the compact
routed experts. Multi-token prediction (MTP) uses the matching draft to
propose tokens that the main model verifies; the default depth is three.
Thinking can be enabled for reasoning or disabled for direct responses.
The published task comparison below uses thinking off throughout.
Benchmarks at a glance
The complete benchmark report contains all six tested stacks: CIRU v4.4.1, CIRU v5.0, HaloBox, Gufo, AgentionAI AP-Q5_K_XL, and Halogen. It includes all 18 prompt/depth points per stack through 256,000 prior tokens, each requesting 512 output tokens; full-prefix prefill; objective accuracy; resident memory; HumanEval 0โ9 speed; the 15-case hard tool set; and two HermesAgent-20 passes. Download the 108-row speed CSV or numerical results.
Current release versus retained v4.4.1
Same Ryzen AI Max+ 395 / Radeon 8060S, one slot, 262,144 context, F16 attention cache and MTP depth 3. Each request appends 2,048 new tokens to the cached depth, then generates 512. Each row is one observation; PP is prompt tok/s and TG is generation tok/s. TG uses the server's 511 inter-token timing convention.
| Prior depth | v4.4.1 PP | v5.0 PP | v4.4.1 TG | v5.0 TG |
|---|---|---|---|---|
| 64,000 | 716.52 | 902.73 | 42.91 | 55.48 |
| 256,000 | 453.78 | 675.56 | 32.58 | 51.39 |
The full six-stack suite used the preceding v5.0 build. The release build above preserves the measured output token IDs and selected attention positions/scores, but the full quality panel and task suite were not rerun after its final long-context change.
Complete-suite quality, memory and behavior
The objective panel scores 2,048 next-token positions against the same BF16 full-vocabulary reference. Lower KL/PPL is better; higher top-1 agreement is better. Native rows use their actual width-four path. Memory below is loaded GPU-visible unified-memory allocation, not total physical memory. These are different complete weights/runtime stacks.
| Stack | KL โ | Top-1 โ | Tail PPL โ | Ready GPU GiB |
|---|---|---|---|---|
| CIRU v4.4.1 | 0.056117 | 93.750% | 2.05516 | 90.60 |
| CIRU v5.0 | 0.050030 | 94.482% | 2.04505 | 87.20 |
| HaloBox | 0.029182 | 96.143% | 2.03642 | 93.40 |
| Gufo | 0.027966 | 96.436% | 2.03502 | 88.54 |
| AgentionAI AP | 0.029135 | 96.436% | 2.02636 | 99.64 |
| Halogen | 0.085905 | 92.773% | 2.07968 | 34.47โ |
โ Halogen also had 78.16 GiB engine process PSS at 64K. GTT, RSS/PSS and system RAM overlap on unified memory and must not be added. See the resident-memory details. The small quality panel does not establish long-context factual accuracy.
| Stack | HumanEval speed TG | Hard tools /100 | Hermes pass 1 /100 | Hermes pass 2 /100 |
|---|---|---|---|---|
| CIRU v4.4.1 | 60.77 | 67 | 89 | 88 |
| CIRU v5.0 | 55.70 | 80 | 93 | 96 |
| HaloBox | 45.84 | 63 | 90 | 96 |
| Gufo | 58.75 | 73 | 95 | 95 |
| AgentionAI AP | 46.42 | 73 | 91 | 91 |
| Halogen | 69.28 | 70 | 93 | 92 |
HumanEval 0โ9 produced ten normal first responses on every stack and was used for speed, not code pass@1. v5.0 was slower on this short-generation check than v4.4.1. Hermes uses a fixed turn budget, and scores can coexist with incomplete runs; exact completion counts, harness corrections and client-error audits are in the report.
Choose a version
| File | Purpose | Bytes | SHA-256 |
|---|---|---|---|
Qwen3.8-Flash-CIRU-STRIX-IU4.gguf |
Retained v4.4.1 | 79,397,818,720 | c0ea11e4e24d0f909720b6c4e7462aa1e6fbf5e0f6acc796063f2aed4cf46ed0 |
Qwen3.8-Flash-CIRU-STRIX-IU4-v5.0.gguf |
New compact v5.0 | 72,112,333,312 | 745bf5b48ef5bd4ded36b3136cde3d35145b2ec5a3aabf03da0e5a792bc5a27d |
mtp/Qwen3.8-Flash-CIRU-STRIX-IU4-MTP-Q8_0.gguf |
Shared draft | 4,135,893,248 | e6743badef1f2619fcb5addfa4344a2a3368cb75214735117e3af80c70b80642 |
ple/ple.payload.bin |
Required SSD-paged embedding table | 52,429,053,952 | 687fc742efb6888c6cd7cf9c80cb4b1ac8cb4707b9409c206699c43363e239b2 |
The other required embedding files are ple.manifest.json and
ple.scale.bf16; the v5.0 model and installation files are listed in
checksums-v5.0.1.sha256. The v4.4.1 file is retained
unchanged for existing users and can be downloaded alone with its matching
runtime.
Download and run v5.0.1
Download the versioned GGUF, matching MTP draft, the three embedding-table
files, the matching vision projector, and the complete source package. Existing users can reuse the same
verified MTP and ple/ files.
python3 -m venv .venv-hf
.venv-hf/bin/pip install --upgrade huggingface_hub
.venv-hf/bin/hf download jcbtc/Qwen3.8-Flash-CIRU-STRIX-IU4 \
--revision v5.0.1 \
Qwen3.8-Flash-CIRU-STRIX-IU4-v5.0.gguf \
mtp/Qwen3.8-Flash-CIRU-STRIX-IU4-MTP-Q8_0.gguf \
ple/ple.payload.bin ple/ple.manifest.json ple/ple.scale.bf16 \
vision/mmproj-Qwen3.8-Flash-F16.mmproj \
source-tree-flash-v5.0.1.tar.gz run-server-v5.0.1.sh INSTALL.md checksums-v5.0.1.sha256 \
--local-dir ./flash-model
cd flash-model
sha256sum -c checksums-v5.0.1.sha256
On Ubuntu/Debian, install python3-venv first if needed. Continue with
INSTALL.md for the complete source build and launch steps.
The included source archive supplies the v5.0 inference code; the guide builds
its pinned AMD SDK and custom HIP/ROCr dependencies from public sources.
It also covers health checks, chat, function calling, thinking settings,
image input, and common setup errors. No earlier CIRU installation or
unpublished GitHub tag is required.
Use 128 GiB unified memory and fast local NVMe for the measured configuration. The model files take about 120 GiB, with additional space needed for the SDK and build. The default local server uses one slot, 16,384-token batches, 262,144-token context, F16 attention cache, and MTP depth three. Performance was measured on NixOS; the Ubuntu/Debian guide is a source-build route, not a claim of equal measured performance on every distribution.
The download includes the matching vision projector, and the v5.0.1 launcher enables it by default. Text benchmark results do not measure image performance. See vision/README.md for the image test scope.
To keep using v4.4.1, follow its retained model card and instructions. Its GGUF, matching runtime packages, original launcher, and checksums remain available. The new model, source archive and launcher use separate filenames.
Historical v4 benchmark record
Earlier 2026 builds of the retained v4 model recorded HumanEval 160/164, HumanEval+ 155/164, ARC-Challenge 1,143/1,172, ToolEval Standard 115/138 and ToolEval Hard 23/30. These were separate local runs with earlier runtime/settings, not reruns of the v5.0 model and not directly comparable to the small thinking-off checks above. The release history preserves v4.1โv4.4.1 changes, v3 comparisons, their caveats and immutable old instructions. Detailed quality history is in BENCHMARKS.md.
Credits
Qwen produced Qwen3.8-Flash-Next and its FP8 source. CIRU builds on llama.cpp/ggml. The Strix Halo runtime draws on the work of pwilkin and the Halogen project. Ported Q5 kernels and the expert prefill architecture use Gufo work; relevant source notices and its MIT terms are retained. The Qwen experimental/MTP work credits Ryan Monsurate; the chunked GDN gate incorporates a HaloBox lead. AMD's ROCm/TheRock projects provide the toolchain. See NOTICE.md and the runtime source notices for provenance. This community release is independent of Qwen and AMD.
- Downloads last month
- 16,307
8-bit