Qwen 3.8 27B ROCmFPX (GGUF for AMD Radeon GPUs)

ROCmFPX Server Qwen 3.8 Project Hardware Vulkan License

This repository contains optimized ROCmFP4 and ROCmFP8 GGUF releases of Qwen 3.8 27B, custom-engineered for AMD Strix Halo (Ryzen AI Max+ 395 / Radeon 8060S) and AMD RDNA4 discrete GPUs (Radeon RX 9070 / 9070 XT).

🛠️ Unified Model Server: github.com/julianmb/rocmfpx-server (Multi-model zoo, hot-swapping, OpenAI API)
📖 Qwen 3.8 Dedicated Deep-Dive: github.com/julianmb/q38rocm (Standalone optimization report, benchmarks & tuning)


📦 Available GGUF Quantizations

File Name Precision File Size Raw Decode MTP Speculative Decode Best For
Qwen3.8-27B-ROCmFP4-FAST.gguf 4.26 bpw 13.55 GiB 14.02 tok/s 🔥 30.56 – 36.04 tok/s Maximum Speed (Gold Standard, Fits 16GB RX 9070 XT!)
Qwen3.8-27B-ROCmFP8.gguf 8.25 bpw 26.25 GiB 7.66 tok/s 18.96 tok/s Maximum Precision (<0.003 PPL loss vs FP16)
Qwen3.8-27B-ROCmFP4-STRIX_LEAN.gguf 4.34 bpw 13.82 GiB 13.90 tok/s 30.00 – 34.00 tok/s FP16 embeddings & norm preservation
Qwen3.8-27B-Q3_K_S.gguf 3.59 bpw 11.40 GiB 16.69 tok/s 20.44 – 26.11 tok/s Fastest unassisted decode
Qwen3.8-27B-Q3_K_M.gguf 3.95 bpw 12.56 GiB 15.15 tok/s 25.00 – 28.50 tok/s Balanced 3-bit medium quantization
Qwen3.8-27B-ROCmFP2.gguf 2.69 bpw 8.56 GiB 12.82 tok/s 17.50 – 19.00 tok/s Ultra-compact memory footprint

🔬 Why ROCmFP4 & ROCmFP8 Improve Performance on AMD Hardware

In auto-regressive decoding, generation speed is physically bound by memory bandwidth:

  • Bandwidth Math: Strix Halo's 256-bit memory controller delivers 190–200 GB/s sustained read bandwidth. Loading an unquantized FP16 model (54.6 GB) caps generation at **4.9 tok/s**.
  • 8-Bit (ROCmFP8 - 26.25 GB): Doubles memory streaming throughput to 18.96 tok/s with MTP speculation while providing lossless precision (<0.003 PPL loss vs FP16). While RDNA executes matrix ALUs in FP16, transferring 8-bit bytes across the bus cuts memory transit time in half.
  • 4-Bit (ROCmFP4 - 13.55 GB): Slashes memory payload by 75.2% vs FP16, raising unassisted streaming decode to 14.02 tok/s and MTP speculative throughput to 30.56 – 36.04 tok/s with ~99% benchmark retention.

🤝 Upstream Engine Attribution

Built using the open-source charlie12345/ROCmFPX toolchain, the pioneering ROCmFPX/ROCmFP4 fork of llama.cpp created by Charlie and the AMD GPU community.


📊 Performance Benchmarks (AMD Ryzen AI Max+ 395)

Measured on AMD Ryzen AI Max+ 395 (40 CU Radeon 8060S @ 2.9 GHz, 128 GB 256-bit LPDDR5X, Linux 7.0, Mesa 26.0 RADV):

Optimization Profile Model Size Unassisted Decode (Measured) MTP Speculative Decode (Measured) Speedup vs Baseline TTFT (Prompt Eval) (Measured)
Stock Q4_K_M (Baseline) 15.92 GiB 12.27 tok/s N/A 1.00× 526.7 ms
ROCmFP4_FAST (This Model) 13.55 GiB 14.02 tok/s N/A 1.14× 468.3 ms
ROCmFP4_FAST + Strict Greedy MTP 13.55 GiB 14.02 tok/s 34.82 tok/s 2.84× 442.8 ms
ROCmFP4_FAST + MTP (n6/p0.60) 13.55 GiB 14.02 tok/s 30.56 – 34.82 tok/s 2.50× – 2.84× 439.4 ms
ROCmFP4_FAST + Deep Spec (n7/p0.35) 13.55 GiB 14.02 tok/s 🔥 36.04 tok/s (JSON/Code) 🔥 2.94× 445.8 ms

💾 Context Scaling & Memory Footprint

Using Asymmetric TurboQuant KV cache (-ctk q8_0 -ctv turbo4):

Context Window Model Weights TurboQuant KV Cache Total RAM Footprint
8K tokens 13.55 GiB 0.62 GiB 14.17 GiB (Fits 16GB RX 9070 XT)
16K tokens 13.55 GiB 1.23 GiB 14.78 GiB (Fits 16GB RX 9070 XT)
32K tokens 13.55 GiB 2.45 GiB 16.00 GiB (Ideal for 64GB Strix Halo)
64K tokens 13.55 GiB 4.90 GiB 18.45 GiB
128K tokens 13.55 GiB 9.80 GiB 23.35 GiB
262K tokens (Full) 13.55 GiB 20.08 GiB 33.63 GiB

📥 Quick Download

# Using official HF CLI
hf download julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF Qwen3.8-27B-ROCmFP4-FAST.gguf --local-dir .

# Or download ROCmFP8 (Zero-loss 8-bit precision)
hf download julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF Qwen3.8-27B-ROCmFP8.gguf --local-dir .

🚀 How to Run

Option 1: Via rocmfpx-server (Recommended Multi-Model Server)

git clone https://github.com/julianmb/rocmfpx-server.git
cd rocmfpx-server
pip install -e .

rocmfpx load qwen38-27b

Option 2: Standalone via q38rocm

git clone https://github.com/julianmb/q38rocm.git
cd q38rocm

./quickstart.sh

🔒 Limitations & Safety

  • Custom Backend: Requires the ROCmFPX toolchain (pinned build: e87d53e (213)).
  • Hardware Target: Optimized for AMD Strix Halo (gfx1151) and AMD RDNA4 (gfx1201).
  • Base Alignment: Inherits base safety characteristics and knowledge capabilities of Qwen 3.8 27B.

📜 License & Attribution

  • Base Model: Qwen 3.8 27B by Alibaba Cloud
  • Quantization & Optimizations: Apache 2.0 License.
  • Community Research: Multi-slot concurrency metrics referenced from MrWidmoreHK and kujetic.
Downloads last month
4,026
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(536)
this model

Evaluation results

  • Peak Speculative Decode Speed on Strix Halo LLM Benchmark Suite
    self-reported
    36.040
  • Strict Lossless Greedy MTP Speed on Strix Halo LLM Benchmark Suite
    self-reported
    34.820
  • Base Unassisted Decode Speed on Strix Halo LLM Benchmark Suite
    self-reported
    14.020
  • Prompt Evaluation Latency (TTFT) on Strix Halo LLM Benchmark Suite
    self-reported
    439.400
  • Effective Bits Per Weight on Strix Halo LLM Benchmark Suite
    self-reported
    4.260