actually a longer context does not take up more embedding params. though it sometimes can depending on the positional embeddings used; if you are using RoPE more or less context doesnt affect param count. Out of the most popular types only learned absolute embeddings will always create more params whereas relative learned embeddings will sometimes create more. most others dont.
Hoglet (Ash) PRO
Hoglet-33
P(doom) <0.1%
AI & ML interests
Open source AI, datasets, parameter efficiency, SLMs, AI for the betterment of humanity. Contact at ash@basicallyai.co
Recent Activity
new activity about 1 hour ago
AxiomicLabs/GPT-X3-150M:WOW! new activity about 3 hours ago
basically-ai/Pebble-10M-Chat:Holy sh, a small language model that is mamab! repliedto Banaxi-Tech's post about 5 hours ago
Well GPT X3 takes the lead. As of right now.
We are now announcing BananaMind 3 π! (not ai for those emoji guys)
All models will use BGA (which is almost just NSA) and our BM3X architecture.
Its sizes will be:
BananaMind 3 Flash Lite, 3M parameters at a context of 8K context.
BananaMind 3 Lite, 10M parameters with 16K context.
BananaMind 3 Flash, 25M Parameters with 16K context.
BananaMind 3 Pro, 50M parameters with 24K context.
BananaMind 3 Ultra, 100M parameters with 32K context.
And lastly, BananaMind 3 Max with 150M parameters and 64K CONTEXT.
I can assure you BananaMind 3 Max WILL beat GPT X3 or match it, we won't release it otherwise. We hope for a 40+ INTELLIGENCE INDEX!
BananaMind 3 may also be partnered with dot labs.
We will cancel BananaMind 2.1 and BananaMind 2 Ultra.
As of the BETU SLM Leaderboard we may need to release it after October 11, im very busy right now (even though we said We will release BETU leaderboard before Oct 11 π)