q6 mtp issues.

#3
by magnade - opened

I'm seeing an issue with the q6 version where MTP isn't a speed up, I see say 1600 token/s with prompt processing with MTP off and with MTP enabled I see 200 ish
I did try the q8 MTP file and it got me 100 ish instead.
I don't see the problem with ornith 1.0 mtp that protolabs did a while back, I do see a minor speed loss on prompt processing of maybe 200 so still starts around 1400.
token gen is impacted also but prompt processing shows it quicker of course.

protoLabsAI org

Thanks for the report β€” and for the 1.0 comparison, that detail is what made this diagnosable.

I couldn't reproduce it on a fully-resident GPU. Q6_K on an RTX PRO 6000, ctx 16K, -ngl 99, flash-attn on:

prompt tokens   arm       pp tok/s   tg tok/s
2548            no-MTP        8059      171.8
2548            +MTP          6793      231.9
10951           no-MTP        7982      168.3
10951           +MTP          7227      272.0

So MTP's real cost to prompt processing is 10-28%, not 8x, and decode speeds up as advertised. (I checked flash-attn off too β€” no change.) That means the files are fine and something in the environment is doing this, so I went looking for what.

I think you're running out of VRAM, and I think our card told you to.

Your symptoms are the signature of a VRAM overflow rather than a compute problem: prompt processing dropping by a multiple instead of a percentage, token gen dragged down too, and β€” the tell β€” Q8_0 being worse than Q6_K (100 vs 200). If MTP were inherently slow the bigger quant wouldn't be disproportionately worse. On Windows especially, the driver spills the overflow into system RAM silently instead of failing to load, so you get a server that works and is several times slower.

Here's what MTP actually costs, measured (Q6_K, ctx 4096):

configuration                 resident VRAM
trunk only                        7.5 GB
trunk + MTP                       8.2 GB   (+0.7)
trunk + mmproj                    8.6 GB   (+1.1)
trunk + mmproj + MTP              9.3 GB   (+1.9)

And this is the part that explains why 1.0 was fine for you: the 1.0 and 1.5 trunk files are byte-identical in size (Q6_K 7.04 GB, Q8_0 9.11 GB). The difference is that 1.5 is a vision model and 1.0 isn't β€” so 1.5 has an mmproj, and the run example on our card loaded it unconditionally, even for pure text. Following our own instructions puts you at 9.3 GB where the same rung on 1.0 sat at 7.5 GB. That's enough to tip a 10 GB card over, and Q8_0 (~11.4 GB) tips harder.

That was a genuine documentation bug on our end. The card is now fixed: the run example is text-only, the mmproj is opt-in and labelled with its cost, and there's a VRAM table plus prompt-processing benchmarks (which we'd never published β€” we only ever benchmarked decode, which is how this reached you).

Worth trying, in order:

  1. Drop --mmproj if you're not sending images. Frees ~1.1 GB, costs you nothing on text.
  2. Use IQ4_XS instead of Q6_K β€” 1.6 GB smaller, and it actually benchmarks faster than Q4_K_M with a bigger MTP gain.
  3. Lower --ctx-size.

If it's still slow after dropping the mmproj, I'd like to keep digging β€” could you post your GPU + VRAM, and the first ~20 lines of llama-server startup (the tensor-offload and KV buffer lines)? That'll show directly whether anything is landing on the host, and if it isn't then my theory is wrong and I'll take another run at it.

I'll test the mmproj item tomorrow and report back.
my card is a radeon 9060xt with 16gb vram, using vulkan backend in linux.
currently seeing 12.6gb ram used, guess I could try -fit off also as that does help with dealing with say 35b models that take up to much vram

protoLabsAI org

That rules out the memory starved theory. Currently trying with Vulkan to see if we can reproduce it there.

protoLabsAI org
β€’
edited Aug 23

Scratch the VRAM overflow theory β€” 16 GB with 12.6 GB used isn't starved, and that framing was Windows-specific anyway. Your -fit off hunch looks right though.

MTP inverts as soon as any layer sits on the CPU. Q6_K, ctx 8192, -fit off, explicit -ngl:

-ngl      no-MTP pp   no-MTP tg   +MTP pp   +MTP tg    MTP verdict
34 / 99        8543       171.4      6537     235.1    +37%  WIN
33             5892        77.3      5438      56.0    -28%  LOSS
28             3066        20.2      2846      12.1    -40%  LOSS
24             2150        13.9      2044       8.2    -41%  LOSS

34 layers total (32 trunk + nextn + output). At 33 β€” one short β€” decode is down 55% and MTP has gone from +37% to -28%. Speculation drafts and verifies every step, so it crosses the host boundary far more often than plain decode does.

--fit is on by default and adjusts anything you left unset, -ngl included, to fit device memory with a 1 GB margin. Turning MTP on raises its estimate, so it can drop the layers MTP needs. 12.6 GB used doesn't rule this out β€” --fit stays under budget by moving layers to host. It also fits 1.0 working and 1.5 not: 1.5 is the one with an mmproj, so it sits closer to the line.

The number worth grabbing tomorrow is the offloaded X/Y layers to GPU line at startup. If it isn't 34/34, that's your cause.

  1. -fit off with an explicit -ngl 99 β€” -fit off alone won't do it if -ngl is unset
  2. drop --mmproj unless you're sending images (-1.1 GB)
  3. IQ4_XS if it's still tight β€” 1.6 GB smaller than Q6_K, and faster

One caveat: I didn't reproduce your 8x prompt-processing drop. Partial offload costs me 2-3x, not 8x. If you get to 34/34 and it's still bad, post the startup log β€” beyond that it's RDNA4/RADV territory and I'd be guessing. I did rebuild with the Vulkan backend to rule out the backend itself, which came back clean, but that was Vulkan on NVIDIA, not your driver.

Card's updated with the table and the --fit warning.

fit = off alone fixed the issue, seeing 14.4gb of vram used now
prompt processing is starting ~1600 and tg i saw some peaks near 60
here is what my router config looks like for anyone else that happens to stumble across this thread and wants a working config for a 9060xt

[Ornith-1.5-9B-MTP-Q6]
hf = protoLabsAI/Ornith-1.5-9B-MTP-GGUF:Q6_K
cache-type-k = q8_0
cache-type-v = q8_0
reasoning-preserve = 1
spec-type = draft-mtp,ngram-mod
spec-draft-n-max = 3
spec-draft-p-min = 0.80
image-min-tokens = 1024
ctx-size = 262144
fitc = 262144
fit = off```

also thank you for the help and pointing out the mmproj coming along and confusing me.

protoLabsAI org

Awesome, I'm glad that worked for you. Happy to help

Sign up or log in to comment