Yes. We were wrong. And we are honest about it. Read this to see us explaining what went wrong.
So, two days ago, we wrote this blog showing almost unbelievable performance on our new GatedDeltaNet 5M model. This model actually wasn't as good as we thought. It was much worse.
The acc vs acc_norm trap
As always, we were evaluating the model against our typical LM-Eval tasks (PIQA, HellaSwag, ARC-Easy and ARC-Challenge) and it went pretty straight-forward.
But we accidently forgot to implement using acc_norm into the benchmark script and blindly took acc (with out normalization ðŸ˜).
The new model and the correct results
We do not have the weights for the ~300M tokens model anymore. Sorry. But we retrained and improved the model - now on 5B high-quality webdata tokens. Here are the REAL results, this time with acc_norm!
| Model | Train tokens | ARC-Easy | ARC-Challenge | HellaSwag | PIQA |
|---|---|---|---|---|---|
| Supra-5M-GatedDeltaNet (CURRENT ONE) | 5B | 33.59% | 23.21% | 26.74% | 52.88% |
| fromziro/Qana-mini-5M | ~21B | 34.97% | 23.21% | 27.60% | 57.18% |
| AxiomicLabs/GPT-S2-5M | ~75B | 33.92% | 22.87% | 27.87% | 57.56% |
| User01110/CMA-8M | ~21B | 35.35% | 23.29% | 28.19% | 58.22% |
That's the whole truth. Now with acc_norm.
The new model
As mentioned above, we have trained the model entirely new from scratch - this time on 5B tokens.
What this means for the soon upcoming Supra3
This architecture changes a lot and we'll take a deeper look at it to maybe even implement it into our future work - maybe even into Supra3!
Stay tuned!
SupraLabs_