Thank you! ๐น
...you really made my day! Will give it a try ASAP... โ๏ธ๐ค
Hope it works well for you! Lots of experiments and changes in this one
Thank you, you've done a great job! Can we expect MoE anytime soon?
Thank you, you've done a great job! Can we expect MoE anytime soon?
I'd second the question/request. Non-MOE's are just so slow unless you can fit it all on the GPU.
Either way, I'll keep an eye on your progress.
Assuming nothing new comes out I think I would try and do the same thing on the MoE as I thought the 26b was pretty solid.
Although the plan at the moment is to take a break for a month or so before continuing.
Initial impressions:
- Instruction-following ability appears to be more surface-level compared to the baseline Gemma 31B Instruct, but it does NOT lose the overall grasp of what's going on (which is good enough for RP in most cases).
- Multi-lingual ability is 100% intact (at least with Eastern European languages), unlike many other fine-tunes that mess it up! Good job on not losing it.
I've tested it with a rather sophisticated 20K-tokens-long prompt that establishes a mix of varied character data (from their life experiences to speech examples to instructions, ending with this particular directive):
Always end your response with a terminal punctuation mark. Use a period, question mark, exclamation point, ellipsis, or combined punctuation as appropriate. Even if you include an emoticon, it must be followed by a closing punctuation mark; for example: >_<.
Original Gemma 31B Instruct follows it rigorously, always using ?/./!/... at the end of each output.
MeroMero v2 seems to be more erratic, ignoring the command sometimes. This raises the question of whether it also skims over some other directives, and to be honest with you, I think it might.
So, in a test case with a difficult past experience between {{char}} and {{user}} - a drama of some sort that has been resolved but not without sacrifices from both sides - the baseline Gemma 31B Instruct writes {{char}} in an "awkward reconciliation" style (this is the exact wording of their current relationship status), while MeroMero v2 sticks to "reconciliation" more, not putting enough effort to show the awkwardness, which results in {{char}} talking in an excited, uplifting way where it's not really expected.
For a practical example, imagine {{char}}/{{user}} sharing a history of meeting IRL but it didn't go well; with MeroMero v2, {{char}} is willing to invite {{user}} back, seeming not to mind their past grudges against {{user}}, even though the prompt also explicitly dictates that the past events do carry some emotional weight and should leave a tangible residue, which the baseline Gemma 31B Instruct certainly does show quite well.
Note: I am not willing to give any kind of verdict, so, no saying the model is good or bad. It's just something to ponder, or perhaps to test for yourself. I agree that it's incredibly difficult to fine-tune this model, and who knows, perhaps this IS the ceiling of what's possible.
Yeah there's definitely a limit to what it can hit IMO (with my resources as a single person doing this as a hobby at least) and I agree this model does run into more logic issues than stock / original meromero. My theory is this is an artifact of the higher diversity / entropy in the model, there's just more chances for it to make a mistake so the same level of intelligence doesn't go as far.
Personal opinion is I think this will generally do better for simpler conversations or where swiping / some steering isn't an issue and the variety is preferred over repeated similar responses. Conversations with lots of rules / lore the v1 MeroMero or another model closer to baseline might handle better.
Just tested it with the original Jinja template from Google, the most recent one... Weirdly, in that same test the model appears to be more consistent with all kinds of emoticons followed by closing punctuation mark correctly. I wonder if Jinja plays a role of significance here, or it's only a matter of more fortunate swipes this time? Either way, we shouldn't be too harsh on it, and I do think v2 handles an abundance of rules well enough, definitely not worse than the other fine-tunes out there (ahem, TheDrummer's Artemis, for example - its earlier iterations were dumbing down G4 quite seriously, which isn't the case with MeroMero v2 at all).
IMO it's safe to be recommended even for complex scenarios. It just happens to give a slightly different flavour to them.
yeah it is pretty solid
thanks for making it <3
@AutisticPancake Give the newest Unsloth Jinja template a try--they fixed quite a few of Google's bugs... ๐๐จ
I'm currently still A/B-testing 31B models, unfortunately it's not an "exact" science, especially when even a fixed seed doesn't guarantee a deterministic turn... ๐ฐ
@McG-221
Aren't they basically the same, some small tool-calling fix aside? In any case, the original template provided with MeroMero v2 is different from both Google / Unsloth.
I've done a bit more testing and I do believe it's somehow better at instruction-following now. Chat is quite lively and the model doesn't seem to omit any important details.
So, yeah, that "surface-level instruction following" claim might be a total BS coming out of my mouth previously, lmao.
Anyway, as the others mentioned, it IS solid and it's worth giving a try.