Yes. We were wrong. And we are honest about it. Read this to see us explaining what went wrong.

So, two days ago, we wrote this blog showing almost unbelievable performance on our new GatedDeltaNet 5M model. This model actually wasn't as good as we thought. It was much worse.

The acc vs acc_norm trap

As always, we were evaluating the model against our typical LM-Eval tasks (PIQA, HellaSwag, ARC-Easy and ARC-Challenge) and it went pretty straight-forward.
But we accidently forgot to implement using acc_norm into the benchmark script and blindly took acc (with out normalization 😭).

The new model and the correct results

We do not have the weights for the ~300M tokens model anymore. Sorry. But we retrained and improved the model - now on 5B high-quality webdata tokens. Here are the REAL results, this time with acc_norm!

ModelTrain tokensARC-EasyARC-ChallengeHellaSwagPIQA
Supra-5M-GatedDeltaNet (CURRENT ONE)5B33.59%23.21%26.74%52.88%
fromziro/Qana-mini-5M~21B34.97%23.21%27.60%57.18%
AxiomicLabs/GPT-S2-5M~75B33.92%22.87%27.87%57.56%
User01110/CMA-8M~21B35.35%23.29%28.19%58.22%

That's the whole truth. Now with acc_norm.

The new model

As mentioned above, we have trained the model entirely new from scratch - this time on 5B tokens.

What this means for the soon upcoming Supra3

This architecture changes a lot and we'll take a deeper look at it to maybe even implement it into our future work - maybe even into Supra3!

Stay tuned!

#research #small-model #gated-delta-net #GDN #we-were-wrong #crying-out-loud-about-acc-norm