Q3 quants generation misfire [solved]

#1
by Nesaliti - opened

Occasionaly and seemingly on random (once in 10 tokens or once in 5000) the model seems to generate gibberish/loose whitespace/chineese for one or few tokens before returning to sane. Couldn't track it down to any of sampling settings to avoid (ST, Text completion). Tried bartowski's IQ3 quant, mradermancher's Q3_K_S and IQ3_S, all three have the same issue.
Are Q3 quants just too broken for modern LLMs to use?

If anyone got consistent results on Q3 quants without any issue on a long run, please share your settings.

Yes I've experienced the same thing. Tried Bart's IQ3_M quant along with Mradermacher's IQ3_S quant and got the same results on both.

Yes I've experienced the same thing. Tried Bart's IQ3_M quant along with Mradermacher's IQ3_S quant and got the same results on both.

Well, I messed around more rn. Dropped the temperature to 0.5 which is insanely low in my opinion (with everything else disabled save for reppen of 1.05 and minp 0.05). The results are consistent, didn't get a single misfire in 10k tokens (so far) but it's... I dunno, always confused?... as well. How's that a considerable downgrade in intelligence from the behemoth on mistral v3 is a mystery.

The prose is interesting though, although very 'slopey'.

With a lot of added empty statements and speech breaks to emphasise-

To emphasise nothing.
Just like that.

Update: Nope, got a misfire again despite low temp - 'Αγγλish'.

Haha yeah, I've noticed those too. I'm going to try bumping up min-p (to 0.1 - so 10%) and keeping top p at 1, top k at 0 and dropping temp down. See if it helps.

I still really like it despite its quirks though!

Oh and I noticed Bartowski's IQ quants are bigger by a few GB than mradermacher. I don't know why. That's why I decided to try one of his out, to see if there's any difference.

Haha yeah, I've noticed those too. I'm going to try bumping up min-p (to 0.1 - so 10%) and keeping top p at 1, top k at 0 and dropping temp down. See if it helps.

I still really like it despite its quirks though!

Okay, I think I found what works fine. Temp 0.3, minp 0.05, reppen 1.05 5196, DRY 0.9 1.75 4 8000, Mirostat 2 3 0.1. Top k/p disabled.
Mirostat visibly affected the output, making it more composed. I hope it also will help to avoid looping.

Model starting to lose it a bit at 32k. 4_K_L. Also bartowski. Behemoth v2 it ain't. But vision on v2 is hit/miss and no tools.

I have no chinese characters except it forgets the plot a little bit, in general, and starts to switch who is who. Maybe this is medium 3.5 issue? It has modern coding/agentic model stink.

I don't want to be the guy looking a gift horse in the mouth but something is.. not quite right. Oh and the reasoning works just fine, its no different than the original. I don't think we're sampling our way out of it.

With last addition of direct I\O in koboldcpp I could save considerable memory of about 15 GB so I could mess around with a bit bigger quants, for many hours by now. Bartowski's Q3_K_S and Q3_K_M (not i-quants) have shown stable work, I stopped with K_M. At temp 0.82 and Mirostat 3.5 I had no problems at all, prose is suddenly fine, no misfires, no stuttering in character's speech or empty looping to fill space in responses. It's geniunely pleasant to work with. I don't have long chat history created with this quant yet (looking up to 24k context) and didn't touch different situations with different starting character and lore prompts, but so far it's pretty good. Nice, really. Too nice sometimes (positivity bias), but that can be fixed via prompt afaik.

Oh, and I forbid a few tokens, namely:
—
Élections
Descart

Don't know if those three are affecting those quants, I just had them pretty regularly with i-quants before.

Nesaliti changed discussion title from Q3 quants generation misfire to Q3 quants generation misfire [solved]
Nesaliti changed discussion status to closed
TheDrummer changed discussion status to open

Sign up or log in to comment